Key takeaways
- The original clue-only condition was the highest pooled Nigerian condition.
- No preloaded model-constructed summary length reproduced fuller-material accuracy.
- Class imbalance and repeated model–case observations constrain inference.
What was tested & why
Cross-stage analysis
The chapter brings fourteen six-model conditions together while separating No Clue and With Clue comparisons. Cases and models recur across conditions; pooled totals should not be read as independent observations.
Source: Thesis Chapter 10, Tables 10.1–10.13 · source-reported unless otherwise noted.
Results / visual evidence
Observed pooled accuracy by material configuration
Thesis Table 8.4 and Figure 8.1 · points mix experimental stages; this is not a controlled causal curve. Full materials average two runs.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
Every model across all ten six-model No Clue conditions
| Model | Orig. | Gen. Mat. | GLMRS sim. | GLMP a | GLMP b | RS-A | RS-B | RS-C | RS-D | RS-E | Mean | Range |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ChatGPT | 70 | 80 | 80 | 90 | 80 | 75 | 90 | 85 | 85 | 70 | 80.5 | 20 |
| Gemini | 60 | 75 | 95 | 75 | 80 | 55 | 65 | 75 | 50 | 65 | 69.5 | 45 |
| Claude | 70 | 70 | 75 | 75 | 85 | 75 | 80 | 85 | 85 | 85 | 78.5 | 15 |
| Grok | 80 | 65 | 45 | 75 | 45 | 45 | 50 | 45 | 40 | 40 | 53 | 40 |
| DeepSeek | 50 | 75 | 65 | 50 | 90 | 80 | 70 | 70 | 75 | 70 | 69.5 | 40 |
| Perplexity | 75 | 70 | 75 | 85 | 65 | 65 | 40 | 45 | 60 | 55 | 63.5 | 45 |
| Pooled | 67.5 | 72.5 | 72.5 | 75 | 74.2 | 65.8 | 65.8 | 67.5 | 65.8 | 64.2 | 69.1 | 10.8 |
Accuracy % over 20 cases. RS-A…E = relational summaries at 50–100, 100–150, 150–300, 300–500, 500–750 words. Pooled range (10.8 points) was far narrower than any model’s range (15–45 points).
Every model across all four With Clue conditions
| Model | Orig. | Gen. Mat. sim. | GLMP a | GLMP b | Mean | Range |
|---|---|---|---|---|---|---|
| ChatGPT | 100 | 100 | 85 | 90 | 93.8 | 15 |
| Gemini | 95 | 70 | 75 | 85 | 81.2 | 25 |
| Claude | 85 | 70 | 85 | 85 | 81.2 | 15 |
| Grok | 80 | 80 | 100 | 45 | 76.2 | 55 |
| DeepSeek | 65 | 75 | 50 | 85 | 68.8 | 35 |
| Perplexity | 80 | 60 | 85 | 80 | 76.2 | 25 |
| Pooled | 84.2 | 75.8 | 80 | 78.3 | 79.6 | 8.3 |
Accuracy % over 20 cases. Grok run-a and Perplexity run-b With Clue outputs carry the provenance flags described in thesis §7.9.
Effect of the case-specific clue by model across the four paired conditions (pp)
| Model | Orig. | Gen. Mat. | GLMP a | GLMP b | Mean |
|---|---|---|---|---|---|
| ChatGPT | +30 | +20 | -5 | +10 | +13.8 |
| Gemini | +35 | -5 | 0 | +5 | +8.8 |
| Claude | +15 | 0 | +10 | 0 | +6.2 |
| Grok | 0 | +15 | +25 | 0 | +10 |
| DeepSeek | +15 | 0 | 0 | -5 | +2.5 |
| Perplexity | +5 | -10 | 0 | +15 | +2.5 |
| Pooled | +16.7 | +3.3 | +5 | +4.2 | +7.3 |
With Clue minus paired No Clue, percentage points; model-level effects rest on 20 cases.
Correct predictions by Nigerian case
Each cell is the number correct out of six models. Select a cell to see the condition and count.
| Case | Ground truth | a NC | a WC | b NC | b WC | A | B | C | D | E |
|---|---|---|---|---|---|---|---|---|---|---|
| NG 001Damisa v. U.B.A. | Dism. | |||||||||
| NG 002Super Ceramics v. H.E.P. Eng. | Dism. | |||||||||
| NG 003Moore Associates v. Exphar | Dism. | |||||||||
| NG 004F.H.A. v. Oyedeji | Dism. | |||||||||
| NG 005N.Y.S.C. v. Ukachukwu | Dism. | |||||||||
| NG 006Akaolisa v. Okuma | Allow. | |||||||||
| NG 007Skye Bank v. Adegun | Dism. | |||||||||
| NG 008Heritage Bank v. Bentworth | Dism. | |||||||||
| NG 009Ethiopian Airlines v. Polaris | Dism. | |||||||||
| NG 010Austin Laz v. GTBank | Dism. | |||||||||
| NG 011Glenyork v. Panalpina | Allow. | |||||||||
| NG 012Olaniran v. Adebayo | Dism. | |||||||||
| NG 013ACMEL v. First Bank | Allow. | |||||||||
| NG 014Total E&P v. Okwu | Allow. | |||||||||
| NG 015A.B.C. Transport v. Omotoye | Dism. | |||||||||
| NG 016Barewa Pharm. v. F.R.N. | Dism. | |||||||||
| NG 017Standard Chartered v. Ameh | Allow. | |||||||||
| NG 018BPS Eng. v. F.R.M.A. | Dism. | |||||||||
| NG 019Omni Products v. Union Bank | Allow. | |||||||||
| NG 020Atiba Iyalamu v. Suberu | Dism. |
a/b = preloaded full-material runs; NC/WC = No Clue/With Clue; A–E = five summary lengths. Source: Thesis Table 10.13. Darker cells mean more correct predictions.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Pooled accuracy across all fourteen six-model conditions (120 observations each, majority-class reference = 84/120 = 70.0%)
Condition Input Loading Correct Accuracy vs 70%
No Clue (original baseline) Facts only – 81 67.5% −2.5 pp
With Clue (original baseline) Facts + clue – 101 84.2% +14.2 pp
General Materials No Clue Fuller materials Simultaneous 87 72.5% +2.5 pp
General Fuller materials + clue Simultaneous 91 75.8% +5.8 pp
Materials With Clue
GLMRS No Clue Relational summaries Simultaneous 87 72.5% +2.5 pp
(150–300 w.)
GLMP a No Clue Fuller materials Preloaded 90 75.0% +5.0 pp
GLMP a With Clue Fuller materials + clue Preloaded 96 80.0% +10.0 pp
GLMP b No Clue Fuller materials (rerun) Preloaded 89 74.2% +4.2 pp
GLMP b With Clue Fuller materials + clue Preloaded 94 78.3% +8.3 pp
(rerun)
GLMRS-A No Clue Relational summaries, Preloaded 79 65.8% −4.2 pp
50–100 w.
GLMRS-B No Clue Relational summaries, Preloaded 79 65.8% −4.2 pp
100–150 w.
GLMRS-C No Clue Relational summaries, Preloaded 81 67.5% −2.5 pp
150–300 w.
GLMRS-D No Clue Relational summaries, Preloaded 79 65.8% −4.2 pp
300–500 w.
GLMRS-E No Clue Relational summaries, Preloaded 77 64.2% −5.8 pp
500–750 w.Thesis Table 10.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model-level totals across the first-stage (five conditions), second- stage (seven conditions) and third-stage (two conditions) experiments
Model First stage Second Third All (/280) Preloaded Mean Mean
(/100) stage stage (/40) range conf., 2nd conf., 3rd
(/140) (/20) stage stage
ChatGPT 86 (86.0%) 117 (83.6%) 33 (82.5%) 236 (84.3%) 14–18 69.7% 57.6%
Gemini 79 (79.0%) 101 (72.1%) 24 (60.0%) 204 (72.9%) 10–17 86.0% 84.6%
Claude 74 (74.0%) 117 (83.6%) 31 (77.5%) 222 (79.3%) 15–17 65.3% 63.7%
Grok 70 (70.0%) 78 (55.7%) 19 (47.5%) 167 (59.6%) 8–20 68.4% 58.8%
DeepSeek 66 (66.0%) 98 (70.0%) 30 (75.0%) 194 (69.3%) 10–18 72.8% 68.9%
Perplexity 72 (72.0%) 95 (67.9%) 21 (52.5%) 188 (67.1%) 8–17 67.3% 65.8%
Pooled 447 606 158 1211 – – –
(74.5%) (72.1%) (65.8%) (72.1%)
Preloaded range is the lowest and highest number correct (out of 20) across the nine preloaded
six-model conditions. Mean confidence is the mean of the model’s condition-level mean
confidence.Thesis Table 10.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled comparison of all ten six-model No Clue conditions (120 observations each)
vs vs
Rank Condition Load Correct Accuracy 95% CI∗
baseline 70%
5 No Clue (original baseline) – 81 67.5% 58.7–75.2 – −2.5 pp
3 General Materials No Clue Sim. 87 72.5% 63.9–79.7 +5.0 pp +2.5 pp
3 GLMRS No Clue Sim. 87 72.5% 63.9–79.7 +5.0 pp +2.5 pp
1 GLMP a No Clue Pre. 90 75.0% 66.6–81.9 +7.5 pp +5.0 pp
2 GLMP b No Clue Pre. 89 74.2% 65.7–81.2 +6.7 pp +4.2 pp
7 GLMRS-A No Clue Pre. 79 65.8% 57.0–73.7 −1.7 pp −4.2 pp
7 GLMRS-B No Clue Pre. 79 65.8% 57.0–73.7 −1.7 pp −4.2 pp
5 GLMRS-C No Clue Pre. 81 67.5% 58.7–75.2 0.0 pp −2.5 pp
7 GLMRS-D No Clue Pre. 79 65.8% 57.0–73.7 −1.7 pp −4.2 pp
10 GLMRS-E No Clue Pre. 77 64.2% 55.3–72.2 −3.3 pp −5.8 pp
∗
Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). “vs baseline” is the difference from the original condition of the
same family, and “vs 70%” is the difference from the always-Dismissed reference. Rank 1 is the
highest pooled accuracy, and ties share a rank. Load: Sim. = simultaneous loading, Pre. =
preloaded.Thesis Table 10.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
No Clue conditions grouped by input type (exploratory pooling of related experiments)
Input type Conditions Correct Accuracy vs facts-only
Facts only (original No Clue) 1 81/120 67.5% –
Fuller materials (simultaneous, preloaded a, 3 266/360 73.9% +6.4 pp
preloaded b)
Relational summaries (first-stage GLMRS, 6 482/720 66.9% −0.6 pp
GLMRS-A to -E)
Preloaded fuller materials only (runs a, b) 2 179/240 74.6% +7.1 pp
Preloaded relational summaries only (A, B, 5 395/600 65.8% −1.7 pp
C, D, E)
Groups combine separate experiments (and, for the preloaded runs, reruns of the same
experiment) and are shown only to summarise direction. They are not independent samples.Thesis Table 10.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy (%) across all ten six-model No Clue conditions
No Gen.
GLMP GLMP
Model Clue Mat. Mean Rng.
GLMRS a b GLMRSGLMRSGLMRSGLMRSGLMRS
(orig.) (sim.)
(sim.) -A -B -C -D -E
ChatGPT 70.0 80.0 80.0 90.0 80.0 75.0 90.0 85.0 85.0 70.0 80.5 20
Gemini 60.0 75.0 95.0 75.0 80.0 55.0 65.0 75.0 50.0 65.0 69.5 45
Claude 70.0 70.0 75.0 75.0 85.0 75.0 80.0 85.0 85.0 85.0 78.5 15
Grok 80.0 65.0 45.0 75.0 45.0 45.0 50.0 45.0 40.0 40.0 53.0 40
DeepSeek 50.0 75.0 65.0 50.0 90.0 80.0 70.0 70.0 75.0 70.0 69.5 40
Perplexity 75.0 70.0 75.0 85.0 65.0 65.0 40.0 45.0 60.0 55.0 63.5 45
Pooled 67.5 72.5 72.5 75.0 74.2 65.8 65.8 67.5 65.8 64.2 69.1 10.8
Entries are accuracy (%) over 20 cases per model. Mean and Range (percentage points) are
taken across the conditions shown. (orig.) = original baseline, Gen. Mat. (sim.) = first-stage
General Materials condition with simultaneous loading, unlabelled = preloaded (GLMP a,
GLMP b and GLMRS-C to -E: second stage, GLMRS-A and -B: third stage).Thesis Table 10.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Change from each model’s original No Clue accuracy (percentage points)
Gen. Above/
GLMP GLMP Mean
Model Mat. equal/
a b GLMRSGLMRSGLMRSGLMRSGLMRS chg. below
(sim.) GLMRS
(sim.) -A -B -C -D -E
ChatGPT +10.0 +10.0 +20.0 +10.0 +5.0 +20.0 +15.0 +15.0 0.0 +11.7 8 / 1 / 0
Gemini +15.0 +35.0 +15.0 +20.0 −5.0 +5.0 +15.0 −10.0 +5.0 +10.6 7 / 0 / 2
Claude 0.0 +5.0 +5.0 +15.0 +5.0 +10.0 +15.0 +15.0 +15.0 +9.4 8 / 1 / 0
Grok −15.0 −35.0 −5.0 −35.0 −35.0 −30.0 −35.0 −40.0 −40.0 −30.0 0 / 0 / 9
DeepSeek +25.0 +15.0 0.0 +40.0 +30.0 +20.0 +20.0 +25.0 +20.0 +21.7 8 / 1 / 0
Perplexity −5.0 0.0 +10.0 −10.0 −10.0 −35.0 −30.0 −15.0 −20.0 −12.8 1 / 1 / 7
Baseline accuracy per model (20 cases): ChatGPT 70.0, Gemini 60.0, Claude 70.0, Grok 80.0,
DeepSeek 50.0, Perplexity 75.0. Model-level changes rest on 20 cases and are noisy.Thesis Table 10.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Reliability indicators for the No Clue conditions (where reported)
Condition Mean Conf. Conf.−Acc. Brier EV SF Incorrect
No Clue (original baseline) n/r n/r n/r n/r 39† 39
General Materials No Clue 81.7% +9.2 n/r n/r n/r 33
GLMRS No Clue 79.8% +7.3 0.1899 +0.3899 33 33
GLMP a No Clue 70.7% −4.3 0.1905 +0.3634 30 30
GLMP b No Clue 68.9% −5.2 0.1819 +0.3642 25 31
GLMRS-A No Clue 66.7% +0.9 0.2424 +0.2073 37 41
GLMRS-B No Clue 66.4% +0.6 0.2219 +0.2297 34 41
GLMRS-C No Clue 70.7% +3.2 0.2222 +0.2562 39 39
GLMRS-D No Clue 71.6% +5.8 0.2452 +0.2261 37 41
GLMRS-E No Clue 70.4% +6.2 0.2386 +0.2063 43 43
n/r = not reported for the Nigerian subset (the original No Clue and With Clue baselines, for
which the primary benchmark reports confidence measures by model across all four
jurisdictions only, and Brier, EV and Silent Failures for the first-stage fuller-material
conditions). † Derived from the primary benchmark: every incorrect No Clue prediction there
was a Silent Failure (134 of 134), and the Nigerian Silent Failure rate of 23.8% over 240
predictions corresponds to 57, leaving 18 under With Clue. Conf.−Acc. is a descriptive gap in
percentage points, not ECE.Thesis Table 10.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled comparison of all four With Clue conditions (120 obser- vations each)
vs vs Paired Clue
Rank Condition Load Correct Accuracy 95% CI∗
baseline 70% No Clue effect
1 With Clue (original – 101 84.2% 76.6–89.6 – +14.2 pp 67.5% +16.7 pp
baseline)
4 General Sim. 91 75.8% 67.4–82.6 −8.3 pp +5.8 pp 72.5% +3.3 pp
Materials With Clue
2 GLMP a With Clue Pre. 96 80.0% 72.0–86.2 −4.2 pp +10.0 pp 75.0% +5.0 pp
3 GLMP b With Clue Pre. 94 78.3% 70.1–84.8 −5.8 pp +8.3 pp 74.2% +4.2 pp
∗
Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). “vs baseline” is the difference from the original condition of the
same family, and “vs 70%” is the difference from the always-Dismissed reference. Rank 1 is the
highest pooled accuracy, and ties share a rank. Load: Sim. = simultaneous loading, Pre. =
preloaded.Thesis Table 10.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy (%) across all four With Clue conditions
Gen.
With Clue
Model Mat. GLMP a GLMP b Mean Rng.
(orig.)
(sim.)
ChatGPT 100.0 100.0 85.0 90.0 93.8 15
Gemini 95.0 70.0 75.0 85.0 81.2 25
Claude 85.0 70.0 85.0 85.0 81.2 15
Grok 80.0 80.0 100.0 45.0 76.2 55
DeepSeek 65.0 75.0 50.0 85.0 68.8 35
Perplexity 80.0 60.0 85.0 80.0 76.2 25
Pooled 84.2 75.8 80.0 78.3 79.6 8.3
Entries are accuracy (%) over 20 cases per model. Mean and Range (percentage points) are
taken across the conditions shown. (orig.) = original baseline, Gen. Mat. (sim.) = first-stage
General Materials condition with simultaneous loading, unlabelled = second-stage preloaded.Thesis Table 10.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Effect of the case-specific clue by model (With Clue minus paired No Clue, percentage points)
Gen.
Original
Model Mat. GLMP a GLMP b Mean Positive / zero / negative
baselines
(sim.)
ChatGPT +30.0 +20.0 −5.0 +10.0 +13.8 3/0/1
Gemini +35.0 −5.0 0.0 +5.0 +8.8 2/1/1
Claude +15.0 0.0 +10.0 0.0 +6.2 2/2/0
Grok 0.0 +15.0 +25.0 0.0 +10.0 2/2/0
DeepSeek +15.0 0.0 0.0 −5.0 +2.5 1/2/1
Perplexity +5.0 −10.0 0.0 +15.0 +2.5 2/1/1
Pooled +16.7 +3.3 +5.0 +4.2 +7.3 –
The Grok run-a and Perplexity run-b With Clue outputs carry the provenance flags described
in the sensitivity analysis. Model-level effects rest on 20 cases.Thesis Table 10.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Reliability indicators for the With Clue conditions (where re- ported)
Condition Mean Conf. Conf.−Acc. Brier EV SF Incorrect
With Clue (original baseline) n/r n/r n/r n/r 18† 19
General Materials With Clue 84.2% +8.4 n/r n/r n/r 29
GLMP a With Clue 77.1% −2.9 0.1515 +0.4861 24 24
GLMP b With Clue 71.9% −6.5 0.1563 +0.4472 20 26
n/r = not reported for the Nigerian subset (the original No Clue and With Clue baselines, for
which the primary benchmark reports confidence measures by model across all four
jurisdictions only, and Brier, EV and Silent Failures for the first-stage fuller-material
conditions). † Derived from the primary benchmark: every incorrect No Clue prediction there
was a Silent Failure (134 of 134), and the Nigerian Silent Failure rate of 23.8% over 240
predictions corresponds to 57, leaving 18 under With Clue. Conf.−Acc. is a descriptive gap in
percentage points, not ECE.Thesis Table 10.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Class-level behaviour in the nine preloaded six-model conditions (Ground Truth: 14 Dismissed and 6 Allowed cases per model)
Condition Dismissed Allowed Balanced Predicted Accuracy −
cases correct cases correct accuracy Allowed 70%
(/84) (/36) (/120)
GLMP a No Clue 72 (85.7%) 18 (50.0%) 67.9% 30 +5.0 pp
GLMP a With Clue 73 (86.9%) 23 (63.9%) 75.4% 34 +10.0 pp
GLMP b No Clue 66 (78.6%) 23 (63.9%) 71.2% 41 +4.2 pp
GLMP b With Clue 69 (82.1%) 25 (69.4%) 75.8% 40 +8.3 pp
GLMRS-A No Clue 62 (73.8%) 17 (47.2%) 60.5% 39 −4.2 pp
GLMRS-B No Clue 59 (70.2%) 20 (55.6%) 62.9% 45 −4.2 pp
GLMRS-C No Clue 63 (75.0%) 18 (50.0%) 62.5% 39 −2.5 pp
GLMRS-D No Clue 60 (71.4%) 19 (52.8%) 62.1% 43 −4.2 pp
GLMRS-E No Clue 61 (72.6%) 16 (44.4%) 58.5% 39 −5.8 pp
All nine 585/756 179/324 66.3% 350/1080 +0.7 pp
(77.4%) (55.2%) (32.4%)
The always-“Appeal Dismissed” strategy scores 14/20 (70.0%) with balanced accuracy 50%.
The Ground Truth contains 6 Allowed cases, i.e. 30.0% of 120 observations.
Models were much better at recognising dismissals than allowances. Over the
nine preloaded conditions they were right on 585/756 Dismissed-case observations
(77.4%) but on only 179/324 Allowed-case observations (55.2%). The GLMRS
conditions had lower accuracy than GLMP on both classes: the Dismissed-case
accuracy was 82.1% for the two GLMP No Clue runs (138/168) against 72.6% for
the five GLMRS conditions (305/420), and the Allowed-case accuracy was 56.9%
(41/72) against 50.0% (90/180). Balanced accuracy was lowest for GLMRS-
E (58.5%), followed by GLMRS-A (60.5%), and highest for the two GLMPThesis Table 10.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Case-level results: number of models (out of 6) predicting the Ground Truth in each preloaded six-model condition
Case Short title GT a NC a WC b NC b WC A B C D E Total %
NG 001 Damisa v. U.B.A. Dism. 6 6 6 6 6 6 6 6 6 54/54 100.0
NG 002 Super Ceramics v. H.E.P. Eng. Dism. 6 4 5 4 6 4 5 4 4 42/54 77.8
NG 003 Moore Associates v. Exphar Dism. 6 6 6 6 5 5 6 5 5 50/54 92.6
NG 004 F.H.A. v. Oyedeji Dism. 6 6 5 5 4 5 5 5 5 46/54 85.2
NG 005 N.Y.S.C. v. Ukachukwu Dism. 5 5 5 5 3 3 4 3 4 37/54 68.5
NG 006 Akaolisa v. Okuma Allow. 3 3 3 2 1 2 2 2 2 20/54 37.0
NG 007 Skye Bank v. Adegun Dism. 5 5 4 5 3 4 4 3 4 37/54 68.5
NG 008 Heritage Bank v. Bentworth Dism. 6 6 5 5 5 5 5 5 5 47/54 87.0
NG 009 Ethiopian Airlines v. Polaris Dism. 1 2 2 2 2 2 2 2 1 16/54 29.6
NG 010 Austin Laz v. GTBank Dism. 6 6 4 4 5 4 3 4 5 41/54 75.9
NG 011 Glenyork v. Panalpina Allow. 1 1 4 4 6 5 5 5 3 34/54 63.0
NG 012 Olaniran v. Adebayo Dism. 6 6 6 6 6 6 6 6 5 53/54 98.1
NG 013 ACMEL v. First Bank Allow. 3 5 5 5 3 6 4 4 3 38/54 70.4
NG 014 Total E&P v. Okwu Allow. 5 5 5 5 4 3 4 4 5 40/54 74.1
NG 015 A.B.C. Transport v. Omotoye Dism. 4 5 5 5 4 4 4 4 4 39/54 72.2
NG 016 Barewa Pharm. v. F.R.N. Dism. 6 6 6 6 6 6 6 6 6 54/54 100.0
NG 017 Standard Chartered v. Ameh Allow. 3 5 2 4 1 2 2 2 1 22/54 40.7
NG 018 BPS Eng. v. F.R.M.A. Dism. 6 6 5 5 5 4 4 4 4 43/54 79.6
NG 019 Omni Products v. Union Bank Allow. 3 4 4 5 2 2 1 2 2 25/54 46.3
NG 020 Atiba Iyalamu v. Suberu Dism. 3 4 2 5 2 1 3 3 3 26/54 48.1
a NC/a WC/b NC/b WC = GLMP run a and rerun b under No Clue and With Clue, A/B =
GLMRS-A/B (third stage, 50–100 and 100–150 words), C/D/E = GLMRS-C/D/E (second
stage, 150–300, 300–500 and 500–750 words). GT = Ground Truth (Dism. = Appeal
Dismissed, Allow. = Appeal Allowed).Thesis Table 10.13 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.