Key takeaways
- Preloaded fuller-material No Clue runs scored 75.0% and 74.2%.
- Six-model identical reruns agreed on about three verdicts in four.
- Unadjusted paired tests are exploratory and clustered observations limit inference.
What was tested & why
Stage 2 · Preloading and reruns
Preloading gave a small descriptive advantage over simultaneous loading. Grok and DeepSeek changed sharply between identical runs, so a single run cannot characterize every condition reliably.
Wilson intervals treat observations as independent and are optimistic; McNemar tests were unadjusted and exploratory.
Source: Thesis Chapter 7 · source-reported unless otherwise noted.
Results / visual evidence
Nigerian extension · selected pooled conditions
Thesis Tables 4.1–4.2, 6.7 and 7.2 · 120 model–case observations per condition; baselines are reused, not recounted.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
Second stage: each model across the seven preloaded conditions
| Model | GLMP a NC | GLMP a WC | GLMP b NC | GLMP b WC | GLMRS-C | GLMRS-D | GLMRS-E |
|---|---|---|---|---|---|---|---|
| ChatGPT | 90% | 85% | 80% | 90% | 85% | 85% | 70% |
| Gemini | 75% | 75% | 80% | 85% | 75% | 50% | 65% |
| Claude | 75% | 85% | 85% | 85% | 85% | 85% | 85% |
| Grok | 75% | 100% | 45% | 45% | 45% | 40% | 40% |
| DeepSeek | 50% | 50% | 90% | 85% | 70% | 75% | 70% |
| Perplexity | 85% | 85% | 65% | 80% | 45% | 60% | 55% |
| Pooled | 75% | 80% | 74.2% | 78.3% | 67.5% | 65.8% | 64.2% |
GLMRS-C/D/E = 150–300, 300–500, 500–750 words per source (No Clue). Identical reruns moved individual models by up to 45 points (Grok WC, 100 → 45; DeepSeek NC, 50 → 90).
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Model accuracy across the seven preloaded conditions: accuracy (number correct out of 20)
GLMP a GLMP a GLMP b GLMP b GLMRS-C GLMRS-D GLMRS-E
Model
No Clue With Clue No Clue With Clue No Clue No Clue No Clue
ChatGPT 90.0% (18) 85.0% (17) 80.0% (16) 90.0% (18) 85.0% (17) 85.0% (17) 70.0% (14)
Gemini 75.0% (15) 75.0% (15) 80.0% (16) 85.0% (17) 75.0% (15) 50.0% (10) 65.0% (13)
Claude 75.0% (15) 85.0% (17) 85.0% (17) 85.0% (17) 85.0% (17) 85.0% (17) 85.0% (17)
Grok 75.0% (15) 100.0% (20) 45.0% (9) 45.0% (9) 45.0% (9) 40.0% (8) 40.0% (8)
DeepSeek 50.0% (10) 50.0% (10) 90.0% (18) 85.0% (17) 70.0% (14) 75.0% (15) 70.0% (14)
Perplexity 85.0% (17) 85.0% (17) 65.0% (13) 80.0% (16) 45.0% (9) 60.0% (12) 55.0% (11)
Pooled 75.0% (90) 80.0% (96) 74.2% (89) 78.3% (94) 67.5% (81) 65.8% (79) 64.2% (77)Thesis Table 7.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled results for the seven preloaded conditions (120 observations each)
Mean Conf.
Condition Correct Accuracy 95% CI∗ Brier EV SF Inc.
conf. −Acc.
GLMP a No Clue 90 75.0% 66.6–81.9 70.7% −4.3 0.1905 +0.3634 30 30
GLMP a With Clue 96 80.0% 72.0–86.2 77.1% −2.9 0.1515 +0.4861 24 24
GLMP b No Clue 89 74.2% 65.7–81.2 68.9% −5.2 0.1819 +0.3642 25 31
GLMP b With Clue 94 78.3% 70.1–84.8 71.9% −6.5 0.1563 +0.4472 20 26
GLMRS-C No Clue 81 67.5% 58.7–75.2 70.7% +3.2 0.2222 +0.2562 39 39
GLMRS-D No Clue 79 65.8% 57.0–73.7 71.6% +5.8 0.2452 +0.2261 37 41
GLMRS-E No Clue 77 64.2% 55.3–72.2 70.4% +6.2 0.2386 +0.2063 43 43
∗
Wilson interval treating the 120 observations as independent. It ignores clustering by case and
by model and is therefore optimistic. Conf.−Acc. is mean confidence minus accuracy in
percentage points (a descriptive gap, not ECE). SF is the number of Silent Failures under the
EV< −0.5 rule.Thesis Table 7.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model-level results for the four GLMP conditions (fuller materials preloaded, run 1 = a, rerun = b)
Condition Model Correct Incorrect Accuracy Mean Conf. Brier EV SF
GLMP a No Clue ChatGPT 18 2 90.0% 67.5% 0.1669 +0.5255 2
Gemini 15 5 75.0% 82.5% 0.1865 +0.4250 5
Claude 15 5 75.0% 67.8% 0.1729 +0.3685 5
Grok 15 5 75.0% 75.8% 0.1843 +0.3855 5
DeepSeek 10 10 50.0% 67.0% 0.2660 +0.0200 10
Perplexity 17 3 85.0% 63.5% 0.1665 +0.4560 3
Pooled 90 30 75.0% 70.7% 0.1905 +0.3634 30
GLMP a With Clue ChatGPT 17 3 85.0% 77.6% 0.1649 +0.5190 3
Gemini 15 5 75.0% 82.5% 0.1865 +0.4250 5
Claude 17 3 85.0% 68.0% 0.1533 +0.4885 3
Grok 20 0 100.0% 93.4% 0.0051 +0.9340 0
DeepSeek 10 10 50.0% 67.0% 0.2660 +0.0200 10
Perplexity 17 3 85.0% 73.9% 0.1334 +0.5300 3
Pooled 96 24 80.0% 77.1% 0.1515 +0.4861 24
GLMP b No Clue ChatGPT 16 4 80.0% 70.7% 0.1789 +0.4215 4
Gemini 16 4 80.0% 87.2% 0.1794 +0.5125 4
Claude 17 3 85.0% 64.8% 0.1687 +0.4585 3
Grok 9 11 45.0% 53.0% 0.2337 −0.0250 5
DeepSeek 18 2 90.0% 78.5% 0.1080 +0.6300 2
Perplexity 13 7 65.0% 59.2% 0.2226 +0.1875 7
Pooled 89 31 74.2% 68.9% 0.1819 +0.3642 25
GLMP b With ClueChatGPT 18 2 90.0% 79.2% 0.1047 +0.6355 2
Gemini 17 3 85.0% 87.2% 0.1344 +0.6075 3
Claude 17 3 85.0% 65.2% 0.1541 +0.4750 3
Grok 9 11 45.0% 53.0% 0.2337 −0.0250 5
DeepSeek 17 3 85.0% 87.2% 0.1131 +0.6275 3
Perplexity 16 4 80.0% 59.2% 0.1976 +0.3625 4
Pooled 94 26 78.3% 71.9% 0.1563 +0.4472 20Thesis Table 7.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model-level results for the three GLMRS conditions (relational summaries preloaded, No Clue)
Condition Model Correct Incorrect Accuracy Mean Conf. Brier EV SF
GLMRS- ChatGPT 17 3 85.0% 62.2% 0.1766 +0.4455 3
C No Clue
Gemini 15 5 75.0% 85.0% 0.2090 +0.4200 5
Claude 17 3 85.0% 65.0% 0.1575 +0.4715 3
Grok 9 11 45.0% 71.5% 0.3206 −0.0725 11
DeepSeek 14 6 70.0% 71.8% 0.1834 +0.3175 6
Perplexity 9 11 45.0% 68.5% 0.2863 −0.0450 11
Pooled 81 39 67.5% 70.7% 0.2222 +0.2562 39
GLMRS- ChatGPT 17 3 85.0% 69.0% 0.1673 +0.4785 3
D No Clue
Gemini 10 10 50.0% 93.0% 0.4218 +0.0150 10
Claude 17 3 85.0% 64.0% 0.1615 +0.4635 3
Grok 8 12 40.0% 65.9% 0.3030 −0.1180 10
DeepSeek 15 5 75.0% 65.5% 0.1710 +0.3600 3
Perplexity 12 8 60.0% 72.2% 0.2466 +0.1575 8
Pooled 79 41 65.8% 71.6% 0.2452 +0.2261 37
GLMRS- ChatGPT 14 6 70.0% 62.0% 0.2200 +0.2530 6
E No Clue
Gemini 13 7 65.0% 84.5% 0.2730 +0.2500 7
Claude 17 3 85.0% 62.3% 0.1760 +0.4430 3
Grok 8 12 40.0% 66.3% 0.3146 −0.1330 12
DeepSeek 14 6 70.0% 72.2% 0.1976 +0.3075 6
Perplexity 11 9 55.0% 74.8% 0.2506 +0.1175 9
Pooled 77 43 64.2% 70.4% 0.2386 +0.2063 43Thesis Table 7.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Model accuracy by relational-summary length (No Clue), with preloaded fuller materials for reference
Model GLMP (mean GLMRS-C GLMRS-D GLMRS-E E − C
of runs a,b) 150–300 300–500 500–750
ChatGPT 85.0% 85.0% 85.0% 70.0% −15.0 pp
Gemini 77.5% 75.0% 50.0% 65.0% −10.0 pp
Claude 80.0% 85.0% 85.0% 85.0% 0.0 pp
Grok 60.0% 45.0% 40.0% 40.0% −5.0 pp
DeepSeek 70.0% 70.0% 75.0% 70.0% 0.0 pp
Perplexity 75.0% 45.0% 60.0% 55.0% +10.0 pp
Pooled 74.6% 67.5% 65.8% 64.2% −3.3 ppThesis Table 7.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled accuracy: simultaneous loading (first stage) versus preload- ing (second stage)
Comparison Simultaneous Preloaded Preloaded Mean of Mean − si-
run a run b runs a,b multaneous
Fuller materials, No Clue 87/120 90/120 89/120 74.6% +2.1 pp
(72.5%) (75.0%) (74.2%)
Fuller materials, With Clue 91/120 96/120 94/120 79.2% +3.3 pp
(75.8%) (80.0%) (78.3%)Thesis Table 7.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Preloaded No Clue comparison by representation and summary length (differences in percentage points)
vs GLMP a vs GLMP b
Condition Correct Accuracy
No Clue No Clue
GLMP a No Clue (fuller materials) 90/120 75.0% – –
GLMP b No Clue (fuller materials, 89/120 74.2% – –
rerun)
GLMRS-C (150–300 words) 81/120 67.5% −7.5 pp −6.7 pp
GLMRS-D (300–500 words) 79/120 65.8% −9.2 pp −8.3 pp
GLMRS-E (500–750 words) 77/120 64.2% −10.8 pp −10.0 ppThesis Table 7.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Effect of adding the case-specific clue under preloading: model- level accuracy and number of verdicts that changed (out of 20)
Run a Run b (rerun) Model No Clue With Clue Changed No Clue With Clue Changed ChatGPT 18/20 17/20 1 16/20 18/20 2 Gemini 15/20 15/20 0 16/20 17/20 1 Claude 15/20 17/20 4 17/20 17/20 2 Grok 15/20 20/20 5 9/20 9/20 0 DeepSeek 10/20 10/20 0 18/20 17/20 1 Perplexity 17/20 17/20 2 13/20 16/20 3 Pooled 90/120 96/120 12 89/120 94/120 9
Thesis Table 7.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Run-to-run stability: verdict agreement between run a and the identical rerun b (out of 20 cases)
No Clue With Clue
Model Same Acc. a Acc. b Same Acc. a Acc. b
verdict verdict
ChatGPT 18/20 90% 80% 17/20 85% 90%
Gemini 15/20 75% 80% 14/20 75% 85%
Claude 18/20 75% 85% 20/20 85% 85%
Grok 12/20 75% 45% 9/20 100% 45%
DeepSeek 12/20 50% 90% 13/20 50% 85%
Perplexity 14/20 85% 65% 17/20 85% 80%
Pooled 89/120 75.0% 74.2% 90/120 80.0% 78.3%
(74.2%) (75.0%)Thesis Table 7.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Mean stated confidence by model and preloaded six-model condi- tion
GLMP a GLMP b
Model
No Clue GLMP a No Clue GLMP bGLMRS-A GLMRS-B GLMRS-C GLMRS-D GLMRS-E
With Clue With Clue No Clue No Clue No Clue No Clue No Clue
ChatGPT 67.5% 77.6% 70.7% 79.2% 54.9% 60.4% 62.2% 69.0% 62.0%
Gemini 82.5% 82.5% 87.2% 87.2% 84.5% 84.8% 85.0% 93.0% 84.5%
Claude 67.8% 68.0% 64.8% 65.2% 63.7% 63.8% 65.0% 64.0% 62.3%
Grok 75.8% 93.4% 53.0% 53.0% 65.3% 52.3% 71.5% 65.9% 66.3%
DeepSeek 67.0% 67.0% 78.5% 87.2% 66.8% 71.0% 71.8% 65.5% 72.2%
Perplexity 63.5% 73.9% 59.2% 59.2% 65.0% 66.5% 68.5% 72.2% 74.8%
Pooled 70.7% 77.1% 68.9% 71.9% 66.7% 66.4% 70.7% 71.6% 70.4%Thesis Table 7.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Sensitivity of the GLMP results (pooled accuracy, and clue effect in percentage points)
Scenario a a b b Clue Clue
No Clue With Clue No Clue With Clue effect a effect b
As scored (six models) 75.0% 80.0% 74.2% 78.3% +5.0 +4.2
Excluding Grok (five models, 75.0% 76.0% 80.0% 85.0% +1.0 +5.0
/100)
Flagged With Clue revisions 75.0% 75.8% 74.2% 75.8% +0.8 +1.7
reverted∗
∗
Grok run a With Clue replaced by its run-a No Clue verdicts, and Perplexity run b
With Clue replaced by its run-b No Clue verdicts (i.e. the recorded “updated” predictions are
not counted).Thesis Table 7.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Exploratory paired comparisons on the same 120 model–case pairs (discordant pairs and unadjusted exact McNemar p)
Comparison (first vs second) First right, Second right, Exact p
second wrong first wrong
GLMP a No Clue vs GLMP b No Clue (rerun) 16 15 1.000
GLMP a With Clue vs GLMP b With Clue 16 14 0.856
(rerun)
GLMP a: No Clue vs With Clue 3 9 0.146
GLMP b: No Clue vs With Clue 2 7 0.180
GLMP a No Clue vs GLMRS-C 21 12 0.163
GLMP a No Clue vs GLMRS-D 23 12 0.090
GLMP a No Clue vs GLMRS-E 20 7 0.019
GLMP b No Clue vs GLMRS-C 14 6 0.115
GLMP b No Clue vs GLMRS-D 16 6 0.052
GLMP b No Clue vs GLMRS-E 19 7 0.029
GLMRS-C vs GLMRS-D 10 8 0.815
GLMRS-D vs GLMRS-E 10 8 0.815
GLMRS-C vs GLMRS-E 13 9 0.523
These tests treat model–case pairs as independent, although they are clustered by case and by
model, and no correction is made for the 13 comparisons. The third-stage comparison of
GLMRS-A with GLMRS-B is reported in Section 8.3. They are reported only to indicate
which differences are clearly within the range of chance variation.Thesis Table 7.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.