Key takeaways
- Original Nigerian No Clue accuracy was 67.5%.
- Simultaneous fuller materials raised No Clue accuracy to 72.5%.
- The original With Clue baseline was 84.2%.
What was tested & why
Stage 1 · Supplying materials
Stage 1 compared fuller materials and relational summaries supplied simultaneously with the prediction task. Representation effects were modest and varied by model. No material-based condition reached the original With Clue benchmark.
Source: Thesis Chapter 6 · source-reported unless otherwise noted.
Results / visual evidence
Nigerian extension · selected pooled conditions
Thesis Tables 4.1–4.2, 6.7 and 7.2 · 120 model–case observations per condition; baselines are reused, not recounted.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
First stage: each model across the five simultaneous-loading conditions
| Model | No Clue | With Clue | Gen. Mat. NC | Gen. Mat. WC | GLMRS NC |
|---|---|---|---|---|---|
| ChatGPT | 70% | 100% | 80% | 100% | 80% |
| Gemini | 60% | 95% | 75% | 70% | 95% |
| Claude | 70% | 85% | 70% | 70% | 75% |
| Grok | 80% | 80% | 65% | 80% | 45% |
| DeepSeek | 50% | 65% | 75% | 75% | 65% |
| Perplexity | 75% | 80% | 70% | 60% | 75% |
| Pooled | 67.5% | 84.2% | 72.5% | 75.8% | 72.5% |
Accuracy over 20 Nigerian cases per model. Pooled GLMRS and fuller materials tie at 72.5%, but Gemini rose 20 points with summaries while Grok fell 20.
GLMRS No Clue: accuracy, calibration and Silent Failures by model
| Model | Accuracy % | Mean conf. % | Brier | EV | Silent Failures /20 |
|---|---|---|---|---|---|
| ChatGPT | 80 | 88.7 | 0.1658 | 0.5365 | 4 |
| Gemini | 95 | 91.6 | 0.0476 | 0.826 | 1 |
| Claude | 75 | 66.4 | 0.1877 | 0.346 | 5 |
| Grok | 45 | 63.6 | 0.2841 | -0.062 | 11 |
| DeepSeek | 65 | 83.7 | 0.2754 | 0.242 | 7 |
| Perplexity | 75 | 85 | 0.1789 | 0.451 | 5 |
| Pooled | 72.5 | 79.8 | 0.1899 | 0.3899 | 33 |
Thesis values; shading applies to the Silent Failure column. Every incorrect GLMRS prediction met the Silent Failure threshold. Grok had 11 of 33 and was the only model with negative mean EV.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Model accuracy across the five first-stage experimental conditions
GLMRS
Model No Clue With Clue
General MaterialsGeneral Materials No Clue
No Clue With Clue
ChatGPT 14/20 (70.0%) 20/20 (100.0%) 16/20 (80.0%) 20/20 (100.0%) 16/20 (80.0%)
Gemini 12/20 (60.0%) 19/20 (95.0%) 15/20 (75.0%) 14/20 (70.0%) 19/20 (95.0%)
Claude 14/20 (70.0%) 17/20 (85.0%) 14/20 (70.0%) 14/20 (70.0%) 15/20 (75.0%)
Grok 16/20 (80.0%) 16/20 (80.0%) 13/20 (65.0%) 16/20 (80.0%) 9/20 (45.0%)
DeepSeek 10/20 (50.0%) 13/20 (65.0%) 15/20 (75.0%) 15/20 (75.0%) 13/20 (65.0%)
Perplexity 15/20 (75.0%) 16/20 (80.0%) 14/20 (70.0%) 12/20 (60.0%) 15/20 (75.0%)
Pooled 81/120 (67.5%) 101/120 (84.2%) 87/120 (72.5%) 91/120 (75.8%) 87/120 (72.5%)Thesis Table 6.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy with fuller general Nigerian legal materials under No Clue
Model Score Accuracy Mean Confidence
ChatGPT 16/20 80.0% 93.30%
Gemini 15/20 75.0% 92.65%
Claude 14/20 70.0% 53.00%
Grok 13/20 65.0% 82.85%
DeepSeek 15/20 75.0% 82.90%
Perplexity 14/20 70.0% 85.75%
Pooled 87/120 72.5% 81.74%Thesis Table 6.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy with fuller general Nigerian legal materials under With Clue
Model Score Accuracy Mean Confidence
ChatGPT 20/20 100.0% 98.25%
Gemini 14/20 70.0% 92.50%
Claude 14/20 70.0% 51.00%
Grok 16/20 80.0% 90.95%
DeepSeek 15/20 75.0% 84.75%
Perplexity 12/20 60.0% 87.80%
Pooled 91/120 75.8% 84.21%Thesis Table 6.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
GLMRS No Clue model-level results
Model Correct Incorrect Accuracy Mean Conf. Brier EV SF ChatGPT 16 4 80.0% 88.7% 0.1658 +0.5365 4 Gemini 19 1 95.0% 91.6% 0.0476 +0.8260 1 Claude 15 5 75.0% 66.4% 0.1877 +0.3460 5 Grok 9 11 45.0% 63.6% 0.2841 −0.0620 11 DeepSeek 13 7 65.0% 83.7% 0.2754 +0.2420 7 Perplexity 15 5 75.0% 85.0% 0.1789 +0.4510 5 Pooled 87 33 72.5% 79.8% 0.1899 +0.3899 33
Thesis Table 6.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled accuracy across general-material No Clue presentations
Condition Correct Total Accuracy
General Nigerian Legal Materials No Clue 87 120 72.5%
GLMRS No Clue 87 120 72.5%Thesis Table 6.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy across the two general-material No Clue presenta- tions
Model Full Materials Relational Summaries
ChatGPT 80.0% 80.0%
Gemini 75.0% 95.0%
Claude 70.0% 75.0%
Grok 65.0% 45.0%
DeepSeek 75.0% 65.0%
Perplexity 70.0% 75.0%
Pooled 72.5% 72.5%Thesis Table 6.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled comparison of the original and added first-stage conditions
Condition Correct Total Accuracy
No Clue 81 120 67.5%
With Clue 101 120 84.2%
General Nigerian Legal Materials No Clue 87 120 72.5%
General Nigerian Legal Materials With Clue 91 120 75.8%
GLMRS No Clue 87 120 72.5%Thesis Table 6.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.