CrossLaw/Putting it together/Cross-stage analysis
15 / 26 · Putting it togetherThesis Chapter 10, Tables 10.1–10.13

Representation matters, but conditions and reruns matter too.

Cross-stage analysis

Key takeaways

  • The original clue-only condition was the highest pooled Nigerian condition.
  • No preloaded model-constructed summary length reproduced fuller-material accuracy.
  • Class imbalance and repeated model–case observations constrain inference.

What was tested & why

Cross-stage analysis

The chapter brings fourteen six-model conditions together while separating No Clue and With Clue comparisons. Cases and models recur across conditions; pooled totals should not be read as independent observations.

Source: Thesis Chapter 10, Tables 10.1–10.13 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Observed pooled accuracy by material configuration

Accuracy— 70% always-Dismissed reference
No added context
67.5%
50–100 words
65.8%
100–150 words
65.8%
150–300 words
67.5%
300–500 words
65.8%
500–750 words
64.2%
Full materials
74.6%

Thesis Table 8.4 and Figure 8.1 · points mix experimental stages; this is not a controlled causal curve. Full materials average two runs.

Model-level comparison

All six models, side by side

Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.

Thesis Table 10.5 · model × all No Clue conditions · source-reported

Every model across all ten six-model No Clue conditions

ModelOrig.Gen. Mat.GLMRS sim.GLMP aGLMP bRS-ARS-BRS-CRS-DRS-EMeanRange
ChatGPT7080809080759085857080.520
Gemini6075957580556575506569.545
Claude7070757585758085858578.515
Grok806545754545504540405340
DeepSeek5075655090807070757069.540
Perplexity7570758565654045605563.545
Pooled67.572.572.57574.265.865.867.565.864.269.110.8

Accuracy % over 20 cases. RS-A…E = relational summaries at 50–100, 100–150, 150–300, 300–500, 500–750 words. Pooled range (10.8 points) was far narrower than any model’s range (15–45 points).

Thesis Table 10.9 · model × all With Clue conditions · source-reported

Every model across all four With Clue conditions

ModelOrig.Gen. Mat. sim.GLMP aGLMP bMeanRange
ChatGPT100100859093.815
Gemini9570758581.225
Claude8570858581.215
Grok80801004576.255
DeepSeek6575508568.835
Perplexity8060858076.225
Pooled84.275.88078.379.68.3

Accuracy % over 20 cases. Grok run-a and Perplexity run-b With Clue outputs carry the provenance flags described in thesis §7.9.

Thesis Table 10.10 · clue effect by model · source-reported

Effect of the case-specific clue by model across the four paired conditions (pp)

ModelOrig.Gen. Mat.GLMP aGLMP bMean
ChatGPT+30+20-5+10+13.8
Gemini+35-50+5+8.8
Claude+150+100+6.2
Grok0+15+250+10
DeepSeek+1500-5+2.5
Perplexity+5-100+15+2.5
Pooled+16.7+3.3+5+4.2+7.3

With Clue minus paired No Clue, percentage points; model-level effects rest on 20 cases.

Table 10.13 · interactive case matrix

Correct predictions by Nigerian case

Each cell is the number correct out of six models. Select a cell to see the condition and count.

CaseGround trutha NCa WCb NCb WCABCDE
NG 001Damisa v. U.B.A.Dism.
NG 002Super Ceramics v. H.E.P. Eng.Dism.
NG 003Moore Associates v. ExpharDism.
NG 004F.H.A. v. OyedejiDism.
NG 005N.Y.S.C. v. UkachukwuDism.
NG 006Akaolisa v. OkumaAllow.
NG 007Skye Bank v. AdegunDism.
NG 008Heritage Bank v. BentworthDism.
NG 009Ethiopian Airlines v. PolarisDism.
NG 010Austin Laz v. GTBankDism.
NG 011Glenyork v. PanalpinaAllow.
NG 012Olaniran v. AdebayoDism.
NG 013ACMEL v. First BankAllow.
NG 014Total E&P v. OkwuAllow.
NG 015A.B.C. Transport v. OmotoyeDism.
NG 016Barewa Pharm. v. F.R.N.Dism.
NG 017Standard Chartered v. AmehAllow.
NG 018BPS Eng. v. F.R.M.A.Dism.
NG 019Omni Products v. Union BankAllow.
NG 020Atiba Iyalamu v. SuberuDism.

a/b = preloaded full-material runs; NC/WC = No Clue/With Clue; A–E = five summary lengths. Source: Thesis Table 10.13. Darker cells mean more correct predictions.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 10.1 · source-reported

Pooled accuracy across all fourteen six-model conditions (120 observations each, majority-class reference = 84/120 = 70.0%)

Condition                       Input                      Loading        Correct   Accuracy   vs 70%

No Clue (original baseline)     Facts only                 –                   81      67.5%    −2.5 pp
With Clue (original baseline)   Facts + clue               –                  101      84.2%   +14.2 pp
General Materials No Clue       Fuller materials           Simultaneous        87      72.5%    +2.5 pp
General                         Fuller materials + clue    Simultaneous        91      75.8%    +5.8 pp
Materials With Clue
GLMRS No Clue                   Relational summaries       Simultaneous        87      72.5%   +2.5 pp
                                (150–300 w.)

GLMP a No Clue                  Fuller materials           Preloaded           90      75.0%    +5.0 pp
GLMP a With Clue                Fuller materials + clue    Preloaded           96      80.0%   +10.0 pp
GLMP b No Clue                  Fuller materials (rerun)   Preloaded           89      74.2%    +4.2 pp
GLMP b With Clue                Fuller materials + clue    Preloaded           94      78.3%    +8.3 pp
                                (rerun)

GLMRS-A No Clue                 Relational summaries,      Preloaded           79      65.8%   −4.2 pp
                                50–100 w.
GLMRS-B No Clue                 Relational summaries,      Preloaded           79      65.8%   −4.2 pp
                                100–150 w.
GLMRS-C No Clue                 Relational summaries,      Preloaded           81      67.5%   −2.5 pp
                                150–300 w.
GLMRS-D No Clue                 Relational summaries,      Preloaded           79      65.8%   −4.2 pp
                                300–500 w.
GLMRS-E No Clue                 Relational summaries,      Preloaded           77      64.2%   −5.8 pp
                                500–750 w.

Thesis Table 10.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.2 · source-reported

Model-level totals across the first-stage (five conditions), second- stage (seven conditions) and third-stage (two conditions) experiments

Model        First stage      Second       Third      All (/280)    Preloaded        Mean         Mean
                 (/100)         stage stage (/40)                       range    conf., 2nd   conf., 3rd
                               (/140)                                    (/20)        stage       stage

ChatGPT       86 (86.0%)   117 (83.6%)   33 (82.5%)   236 (84.3%)        14–18       69.7%        57.6%
Gemini        79 (79.0%)   101 (72.1%)   24 (60.0%)   204 (72.9%)        10–17       86.0%        84.6%
Claude        74 (74.0%)   117 (83.6%)   31 (77.5%)   222 (79.3%)        15–17       65.3%        63.7%
Grok          70 (70.0%)    78 (55.7%)   19 (47.5%)   167 (59.6%)         8–20       68.4%        58.8%
DeepSeek      66 (66.0%)    98 (70.0%)   30 (75.0%)   194 (69.3%)        10–18       72.8%        68.9%
Perplexity    72 (72.0%)    95 (67.9%)   21 (52.5%)   188 (67.1%)         8–17       67.3%        65.8%

Pooled              447          606          158          1211              –           –            –
                (74.5%)      (72.1%)      (65.8%)       (72.1%)

Preloaded range is the lowest and highest number correct (out of 20) across the nine preloaded
    six-model conditions. Mean confidence is the mean of the model’s condition-level mean
                                          confidence.

Thesis Table 10.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.3 · source-reported

Pooled comparison of all ten six-model No Clue conditions (120 observations each)

                                                                                   vs        vs
    Rank Condition                         Load Correct Accuracy 95% CI∗
                                                                                baseline    70%

         5   No Clue (original baseline)   –         81     67.5%   58.7–75.2         –    −2.5 pp
         3   General Materials No Clue     Sim.      87     72.5%   63.9–79.7   +5.0 pp    +2.5 pp
         3   GLMRS No Clue                 Sim.      87     72.5%   63.9–79.7   +5.0 pp    +2.5 pp
         1   GLMP a No Clue                Pre.      90     75.0%   66.6–81.9   +7.5 pp    +5.0 pp
         2   GLMP b No Clue                Pre.      89     74.2%   65.7–81.2   +6.7 pp    +4.2 pp
         7   GLMRS-A No Clue               Pre.      79     65.8%   57.0–73.7   −1.7 pp    −4.2 pp
         7   GLMRS-B No Clue               Pre.      79     65.8%   57.0–73.7   −1.7 pp    −4.2 pp
         5   GLMRS-C No Clue               Pre.      81     67.5%   58.7–75.2    0.0 pp    −2.5 pp
         7   GLMRS-D No Clue               Pre.      79     65.8%   57.0–73.7   −1.7 pp    −4.2 pp
        10   GLMRS-E No Clue               Pre.      77     64.2%   55.3–72.2   −3.3 pp    −5.8 pp
    ∗
     Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). “vs baseline” is the difference from the original condition of the
same family, and “vs 70%” is the difference from the always-Dismissed reference. Rank 1 is the
  highest pooled accuracy, and ties share a rank. Load: Sim. = simultaneous loading, Pre. =
                                           preloaded.

Thesis Table 10.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.4 · source-reported

No Clue conditions grouped by input type (exploratory pooling of related experiments)

   Input type                                        Conditions      Correct     Accuracy    vs facts-only

   Facts only (original No Clue)                               1      81/120        67.5%                –
   Fuller materials (simultaneous, preloaded a,                3     266/360        73.9%          +6.4 pp
   preloaded b)
   Relational summaries (first-stage GLMRS,                    6     482/720        66.9%          −0.6 pp
   GLMRS-A to -E)
     Preloaded fuller materials only (runs a, b)               2     179/240        74.6%          +7.1 pp
     Preloaded relational summaries only (A, B,                5     395/600        65.8%          −1.7 pp
   C, D, E)

    Groups combine separate experiments (and, for the preloaded runs, reruns of the same
 experiment) and are shown only to summarise direction. They are not independent samples.

Thesis Table 10.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.5 · source-reported

Model accuracy (%) across all ten six-model No Clue conditions

              No      Gen.
                                    GLMP GLMP
Model         Clue    Mat.                                             Mean Rng.
                            GLMRS     a    b  GLMRSGLMRSGLMRSGLMRSGLMRS
             (orig.) (sim.)
                             (sim.)             -A   -B   -C   -D   -E

ChatGPT       70.0   80.0    80.0     90.0    80.0      75.0       90.0   85.0     85.0     70.0   80.5   20
Gemini        60.0   75.0    95.0     75.0    80.0      55.0       65.0   75.0     50.0     65.0   69.5   45
Claude        70.0   70.0    75.0     75.0    85.0      75.0       80.0   85.0     85.0     85.0   78.5   15
Grok          80.0   65.0    45.0     75.0    45.0      45.0       50.0   45.0     40.0     40.0   53.0   40
DeepSeek      50.0   75.0    65.0     50.0    90.0      80.0       70.0   70.0     75.0     70.0   69.5   40
Perplexity    75.0   70.0    75.0     85.0    65.0      65.0       40.0   45.0     60.0     55.0   63.5   45

Pooled       67.5    72.5   72.5     75.0     74.2      65.8       65.8   67.5     65.8     64.2   69.1 10.8

 Entries are accuracy (%) over 20 cases per model. Mean and Range (percentage points) are
taken across the conditions shown. (orig.) = original baseline, Gen. Mat. (sim.) = first-stage
  General Materials condition with simultaneous loading, unlabelled = preloaded (GLMP a,
        GLMP b and GLMRS-C to -E: second stage, GLMRS-A and -B: third stage).

Thesis Table 10.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.6 · source-reported

Change from each model’s original No Clue accuracy (percentage points)

              Gen.                                                   Above/
                            GLMP GLMP                          Mean
Model         Mat.                                                   equal/
                              a    b  GLMRSGLMRSGLMRSGLMRSGLMRS chg. below
             (sim.) GLMRS
                     (sim.)             -A   -B   -C   -D   -E

ChatGPT      +10.0   +10.0   +20.0   +10.0   +5.0    +20.0   +15.0   +15.0    0.0    +11.7 8 / 1 / 0
Gemini       +15.0   +35.0   +15.0   +20.0   −5.0    +5.0    +15.0   −10.0   +5.0    +10.6 7 / 0 / 2
Claude        0.0    +5.0    +5.0    +15.0   +5.0    +10.0   +15.0   +15.0   +15.0   +9.4 8 / 1 / 0
Grok         −15.0   −35.0   −5.0    −35.0   −35.0   −30.0   −35.0   −40.0   −40.0   −30.0 0 / 0 / 9
DeepSeek     +25.0   +15.0    0.0    +40.0   +30.0   +20.0   +20.0   +25.0   +20.0   +21.7 8 / 1 / 0
Perplexity   −5.0     0.0    +10.0   −10.0   −10.0   −35.0   −30.0   −15.0   −20.0   −12.8 1 / 1 / 7

Baseline accuracy per model (20 cases): ChatGPT 70.0, Gemini 60.0, Claude 70.0, Grok 80.0,
     DeepSeek 50.0, Perplexity 75.0. Model-level changes rest on 20 cases and are noisy.

Thesis Table 10.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.7 · source-reported

Reliability indicators for the No Clue conditions (where reported)

        Condition                     Mean Conf. Conf.−Acc. Brier         EV SF Incorrect

        No Clue (original baseline)           n/r       n/r      n/r       n/r 39†     39
        General Materials No Clue          81.7%       +9.2      n/r       n/r n/r     33
        GLMRS No Clue                      79.8%       +7.3   0.1899   +0.3899 33      33
        GLMP a No Clue                     70.7%       −4.3   0.1905   +0.3634 30      30
        GLMP b No Clue                     68.9%       −5.2   0.1819   +0.3642 25      31
        GLMRS-A No Clue                    66.7%       +0.9   0.2424   +0.2073 37      41
        GLMRS-B No Clue                    66.4%       +0.6   0.2219   +0.2297 34      41
        GLMRS-C No Clue                    70.7%       +3.2   0.2222   +0.2562 39      39
        GLMRS-D No Clue                    71.6%       +5.8   0.2452   +0.2261 37      41
        GLMRS-E No Clue                    70.4%       +6.2   0.2386   +0.2063 43      43

n/r = not reported for the Nigerian subset (the original No Clue and With Clue baselines, for
     which the primary benchmark reports confidence measures by model across all four
    jurisdictions only, and Brier, EV and Silent Failures for the first-stage fuller-material
conditions). † Derived from the primary benchmark: every incorrect No Clue prediction there
   was a Silent Failure (134 of 134), and the Nigerian Silent Failure rate of 23.8% over 240
predictions corresponds to 57, leaving 18 under With Clue. Conf.−Acc. is a descriptive gap in
                                  percentage points, not ECE.

Thesis Table 10.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.8 · source-reported

Pooled comparison of all four With Clue conditions (120 obser- vations each)

                                                                   vs       vs        Paired   Clue
Rank Condition              Load Correct Accuracy 95% CI∗
                                                                baseline   70%       No Clue   effect

    1 With Clue (original   –         101     84.2% 76.6–89.6         – +14.2 pp       67.5% +16.7 pp
      baseline)
    4 General               Sim.       91     75.8% 67.4–82.6   −8.3 pp    +5.8 pp     72.5%   +3.3 pp
      Materials With Clue
    2 GLMP a With Clue      Pre.       96     80.0% 72.0–86.2   −4.2 pp +10.0 pp       75.0%   +5.0 pp
    3 GLMP b With Clue      Pre.       94     78.3% 70.1–84.8   −5.8 pp +8.3 pp        74.2%   +4.2 pp
    ∗
     Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). “vs baseline” is the difference from the original condition of the
same family, and “vs 70%” is the difference from the always-Dismissed reference. Rank 1 is the
  highest pooled accuracy, and ties share a rank. Load: Sim. = simultaneous loading, Pre. =
                                           preloaded.

Thesis Table 10.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.9 · source-reported

Model accuracy (%) across all four With Clue conditions

                                       Gen.
                With Clue
Model                                  Mat.            GLMP a           GLMP b      Mean Rng.
                 (orig.)
                                      (sim.)

ChatGPT           100.0               100.0             85.0             90.0        93.8   15
Gemini             95.0               70.0               75.0            85.0        81.2   25
Claude             85.0               70.0               85.0            85.0        81.2   15
Grok               80.0                80.0             100.0            45.0        76.2   55
DeepSeek          65.0                 75.0             50.0             85.0        68.8   35
Perplexity         80.0               60.0               85.0            80.0        76.2   25

Pooled             84.2               75.8              80.0             78.3        79.6   8.3

 Entries are accuracy (%) over 20 cases per model. Mean and Range (percentage points) are
taken across the conditions shown. (orig.) = original baseline, Gen. Mat. (sim.) = first-stage
General Materials condition with simultaneous loading, unlabelled = second-stage preloaded.

Thesis Table 10.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.10 · source-reported

Effect of the case-specific clue by model (With Clue minus paired No Clue, percentage points)

                              Gen.
                Original
Model                         Mat.           GLMP a    GLMP b   Mean Positive / zero / negative
                baselines
                             (sim.)

ChatGPT           +30.0      +20.0             −5.0     +10.0   +13.8                  3/0/1
Gemini            +35.0      −5.0               0.0     +5.0     +8.8                  2/1/1
Claude            +15.0       0.0              +10.0     0.0     +6.2                  2/2/0
Grok               0.0       +15.0             +25.0     0.0    +10.0                  2/2/0
DeepSeek          +15.0       0.0               0.0     −5.0     +2.5                  1/2/1
Perplexity        +5.0       −10.0              0.0     +15.0    +2.5                  2/1/1

Pooled            +16.7      +3.3              +5.0     +4.2    +7.3                             –

The Grok run-a and Perplexity run-b With Clue outputs carry the provenance flags described
              in the sensitivity analysis. Model-level effects rest on 20 cases.

Thesis Table 10.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.11 · source-reported

Reliability indicators for the With Clue conditions (where re- ported)

       Condition                       Mean Conf. Conf.−Acc. Brier     EV SF Incorrect

       With Clue (original baseline)           n/r       n/r    n/r     n/r 18†     19
       General Materials With Clue          84.2%       +8.4    n/r     n/r n/r     29
       GLMP a With Clue                     77.1%       −2.9 0.1515 +0.4861 24      24
       GLMP b With Clue                     71.9%       −6.5 0.1563 +0.4472 20      26

n/r = not reported for the Nigerian subset (the original No Clue and With Clue baselines, for
     which the primary benchmark reports confidence measures by model across all four
    jurisdictions only, and Brier, EV and Silent Failures for the first-stage fuller-material
conditions). † Derived from the primary benchmark: every incorrect No Clue prediction there
   was a Silent Failure (134 of 134), and the Nigerian Silent Failure rate of 23.8% over 240
predictions corresponds to 57, leaving 18 under With Clue. Conf.−Acc. is a descriptive gap in
                                  percentage points, not ECE.

Thesis Table 10.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.12 · source-reported

Class-level behaviour in the nine preloaded six-model conditions (Ground Truth: 14 Dismissed and 6 Allowed cases per model)

Condition             Dismissed         Allowed    Balanced      Predicted    Accuracy −
                   cases correct   cases correct   accuracy       Allowed           70%
                           (/84)           (/36)                    (/120)

GLMP a No Clue        72 (85.7%)      18 (50.0%)       67.9%            30         +5.0 pp
GLMP a With Clue      73 (86.9%)      23 (63.9%)       75.4%            34        +10.0 pp
GLMP b No Clue        66 (78.6%)      23 (63.9%)       71.2%            41         +4.2 pp
GLMP b With Clue      69 (82.1%)      25 (69.4%)       75.8%            40         +8.3 pp
GLMRS-A No Clue       62 (73.8%)      17 (47.2%)       60.5%            39         −4.2 pp
GLMRS-B No Clue       59 (70.2%)      20 (55.6%)       62.9%            45         −4.2 pp
GLMRS-C No Clue       63 (75.0%)      18 (50.0%)       62.5%            39         −2.5 pp
GLMRS-D No Clue       60 (71.4%)      19 (52.8%)       62.1%            43         −4.2 pp
GLMRS-E No Clue       61 (72.6%)      16 (44.4%)       58.5%            39         −5.8 pp

All nine               585/756         179/324        66.3%       350/1080        +0.7 pp
                       (77.4%)         (55.2%)                     (32.4%)

 The always-“Appeal Dismissed” strategy scores 14/20 (70.0%) with balanced accuracy 50%.
        The Ground Truth contains 6 Allowed cases, i.e. 30.0% of 120 observations.


Models were much better at recognising dismissals than allowances. Over the
nine preloaded conditions they were right on 585/756 Dismissed-case observations
(77.4%) but on only 179/324 Allowed-case observations (55.2%). The GLMRS
conditions had lower accuracy than GLMP on both classes: the Dismissed-case
accuracy was 82.1% for the two GLMP No Clue runs (138/168) against 72.6% for
the five GLMRS conditions (305/420), and the Allowed-case accuracy was 56.9%
(41/72) against 50.0% (90/180). Balanced accuracy was lowest for GLMRS-
E (58.5%), followed by GLMRS-A (60.5%), and highest for the two GLMP

Thesis Table 10.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 10.13 · source-reported

Case-level results: number of models (out of 6) predicting the Ground Truth in each preloaded six-model condition

  Case   Short title                    GT      a NC a WC b NC b WC A B C D E Total              %

  NG 001 Damisa v. U.B.A.              Dism.     6    6     6    6   6   6   6   6   6   54/54 100.0
  NG 002 Super Ceramics v. H.E.P. Eng. Dism.     6    4     5    4   6   4   5   4   4   42/54 77.8
  NG 003 Moore Associates v. Exphar    Dism.     6    6     6    6   5   5   6   5   5   50/54 92.6
  NG 004 F.H.A. v. Oyedeji             Dism.     6    6     5    5   4   5   5   5   5   46/54 85.2
  NG 005 N.Y.S.C. v. Ukachukwu         Dism.     5    5     5    5   3   3   4   3   4   37/54 68.5
  NG 006 Akaolisa v. Okuma             Allow.    3    3     3    2   1   2   2   2   2   20/54 37.0
  NG 007 Skye Bank v. Adegun           Dism.     5    5     4    5   3   4   4   3   4   37/54 68.5
  NG 008 Heritage Bank v. Bentworth    Dism.     6    6     5    5   5   5   5   5   5   47/54 87.0
  NG 009 Ethiopian Airlines v. Polaris Dism.     1    2     2    2   2   2   2   2   1   16/54 29.6
  NG 010 Austin Laz v. GTBank          Dism.     6    6     4    4   5   4   3   4   5   41/54 75.9
  NG 011 Glenyork v. Panalpina         Allow.    1    1     4    4   6   5   5   5   3   34/54 63.0
  NG 012 Olaniran v. Adebayo           Dism.     6    6     6    6   6   6   6   6   5   53/54 98.1
  NG 013 ACMEL v. First Bank           Allow.    3    5     5    5   3   6   4   4   3   38/54 70.4
  NG 014 Total E&P v. Okwu             Allow.    5    5     5    5   4   3   4   4   5   40/54 74.1
  NG 015 A.B.C. Transport v. Omotoye Dism.       4    5     5    5   4   4   4   4   4   39/54 72.2
  NG 016 Barewa Pharm. v. F.R.N.       Dism.     6    6     6    6   6   6   6   6   6   54/54 100.0
  NG 017 Standard Chartered v. Ameh Allow.       3    5     2    4   1   2   2   2   1   22/54 40.7
  NG 018 BPS Eng. v. F.R.M.A.          Dism.     6    6     5    5   5   4   4   4   4   43/54 79.6
  NG 019 Omni Products v. Union Bank Allow.      3    4     4    5   2   2   1   2   2   25/54 46.3
  NG 020 Atiba Iyalamu v. Suberu       Dism.     3    4     2    5   2   1   3   3   3   26/54 48.1

a NC/a WC/b NC/b WC = GLMP run a and rerun b under No Clue and With Clue, A/B =
  GLMRS-A/B (third stage, 50–100 and 100–150 words), C/D/E = GLMRS-C/D/E (second
    stage, 150–300, 300–500 and 500–750 words). GT = Ground Truth (Dism. = Appeal
                           Dismissed, Allow. = Appeal Allowed).

Thesis Table 10.13 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.