CrossLaw/Part II · Nigeria/Stage 2 · Preloading and reruns
12 / 26 · Part II · NigeriaThesis Chapter 7

Preloading was modestly higher, but reruns moved.

Stage 2 · Preloading and reruns

Key takeaways

  • Preloaded fuller-material No Clue runs scored 75.0% and 74.2%.
  • Six-model identical reruns agreed on about three verdicts in four.
  • Unadjusted paired tests are exploratory and clustered observations limit inference.

What was tested & why

Stage 2 · Preloading and reruns

Preloading gave a small descriptive advantage over simultaneous loading. Grok and DeepSeek changed sharply between identical runs, so a single run cannot characterize every condition reliably.

Evidence boundary

Wilson intervals treat observations as independent and are optimistic; McNemar tests were unadjusted and exploratory.

Source: Thesis Chapter 7 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Nigerian extension · selected pooled conditions

Accuracy— 70% always-Dismissed reference
Original No Clue
67.5%
Simultaneous full materials
72.5%
Preloaded full · run a
75.0%
Preloaded full · run b
74.2%
Original With Clue
84.2%

Thesis Tables 4.1–4.2, 6.7 and 7.2 · 120 model–case observations per condition; baselines are reused, not recounted.

Model-level comparison

All six models, side by side

Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.

Thesis Table 7.1 · model × preloaded condition · source-reported

Second stage: each model across the seven preloaded conditions

ModelGLMP a NCGLMP a WCGLMP b NCGLMP b WCGLMRS-CGLMRS-DGLMRS-E
ChatGPT90%85%80%90%85%85%70%
Gemini75%75%80%85%75%50%65%
Claude75%85%85%85%85%85%85%
Grok75%100%45%45%45%40%40%
DeepSeek50%50%90%85%70%75%70%
Perplexity85%85%65%80%45%60%55%
Pooled75%80%74.2%78.3%67.5%65.8%64.2%

GLMRS-C/D/E = 150–300, 300–500, 500–750 words per source (No Clue). Identical reruns moved individual models by up to 45 points (Grok WC, 100 → 45; DeepSeek NC, 50 → 90).

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 7.1 · source-reported

Model accuracy across the seven preloaded conditions: accuracy (number correct out of 20)

             GLMP a       GLMP a        GLMP b        GLMP b       GLMRS-C GLMRS-D GLMRS-E
Model
             No Clue      With Clue     No Clue       With Clue     No Clue No Clue No Clue

ChatGPT      90.0% (18)   85.0% (17)    80.0% (16)    90.0% (18)   85.0% (17)   85.0% (17)   70.0% (14)
Gemini       75.0% (15)    75.0% (15)   80.0% (16)    85.0% (17)   75.0% (15)   50.0% (10)   65.0% (13)
Claude       75.0% (15)    85.0% (17)   85.0% (17)    85.0% (17)   85.0% (17)   85.0% (17)   85.0% (17)
Grok         75.0% (15)   100.0% (20)    45.0% (9)     45.0% (9)    45.0% (9)    40.0% (8)    40.0% (8)
DeepSeek     50.0% (10)   50.0% (10)    90.0% (18)    85.0% (17)   70.0% (14)   75.0% (15)   70.0% (14)
Perplexity   85.0% (17)    85.0% (17)   65.0% (13)    80.0% (16)    45.0% (9)   60.0% (12)   55.0% (11)

Pooled       75.0% (90) 80.0% (96) 74.2% (89) 78.3% (94) 67.5% (81) 65.8% (79) 64.2% (77)

Thesis Table 7.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.2 · source-reported

Pooled results for the seven preloaded conditions (120 observations each)

                                                     Mean Conf.
       Condition          Correct Accuracy 95% CI∗               Brier       EV SF Inc.
                                                     conf. −Acc.

       GLMP a No Clue          90     75.0% 66.6–81.9 70.7%   −4.3 0.1905 +0.3634   30   30
       GLMP a With Clue        96     80.0% 72.0–86.2 77.1%   −2.9 0.1515 +0.4861   24   24
       GLMP b No Clue          89     74.2% 65.7–81.2 68.9%   −5.2 0.1819 +0.3642   25   31
       GLMP b With Clue        94     78.3% 70.1–84.8 71.9%   −6.5 0.1563 +0.4472   20   26
       GLMRS-C No Clue         81     67.5% 58.7–75.2 70.7%   +3.2 0.2222 +0.2562   39   39
       GLMRS-D No Clue         79     65.8% 57.0–73.7 71.6%   +5.8 0.2452 +0.2261   37   41
       GLMRS-E No Clue         77     64.2% 55.3–72.2 70.4%   +6.2 0.2386 +0.2063   43   43
∗
Wilson interval treating the 120 observations as independent. It ignores clustering by case and
   by model and is therefore optimistic. Conf.−Acc. is mean confidence minus accuracy in
percentage points (a descriptive gap, not ECE). SF is the number of Silent Failures under the
                                       EV< −0.5 rule.

Thesis Table 7.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.3 · source-reported

Model-level results for the four GLMP conditions (fuller materials preloaded, run 1 = a, rerun = b)

  Condition         Model        Correct Incorrect Accuracy Mean Conf.   Brier   EV SF

  GLMP a No Clue    ChatGPT          18        2      90.0%      67.5% 0.1669 +0.5255  2
                    Gemini           15        5      75.0%      82.5% 0.1865 +0.4250  5
                    Claude           15        5      75.0%      67.8% 0.1729 +0.3685  5
                    Grok             15        5      75.0%      75.8% 0.1843 +0.3855  5
                    DeepSeek         10       10      50.0%      67.0% 0.2660 +0.0200 10
                    Perplexity       17        3      85.0%      63.5% 0.1665 +0.4560  3
                    Pooled           90       30     75.0%      70.7% 0.1905 +0.3634 30

  GLMP a With Clue ChatGPT           17        3     85.0%       77.6% 0.1649 +0.5190  3
                   Gemini            15        5     75.0%       82.5% 0.1865 +0.4250  5
                   Claude            17        3     85.0%       68.0% 0.1533 +0.4885  3
                   Grok              20        0    100.0%       93.4% 0.0051 +0.9340  0
                   DeepSeek          10       10     50.0%       67.0% 0.2660 +0.0200 10
                   Perplexity        17        3     85.0%       73.9% 0.1334 +0.5300  3
                   Pooled            96       24    80.0%       77.1% 0.1515 +0.4861 24

  GLMP b No Clue    ChatGPT          16        4      80.0%      70.7% 0.1789 +0.4215 4
                    Gemini           16        4      80.0%      87.2% 0.1794 +0.5125 4
                    Claude           17        3      85.0%      64.8% 0.1687 +0.4585 3
                    Grok              9       11      45.0%      53.0% 0.2337 −0.0250 5
                    DeepSeek         18        2      90.0%      78.5% 0.1080 +0.6300 2
                    Perplexity       13        7      65.0%      59.2% 0.2226 +0.1875 7
                    Pooled           89       31     74.2%      68.9% 0.1819 +0.3642 25

  GLMP b With ClueChatGPT            18        2      90.0%      79.2% 0.1047 +0.6355 2
                  Gemini             17        3      85.0%      87.2% 0.1344 +0.6075 3
                  Claude             17        3      85.0%      65.2% 0.1541 +0.4750 3
                  Grok                9       11      45.0%      53.0% 0.2337 −0.0250 5
                  DeepSeek           17        3      85.0%      87.2% 0.1131 +0.6275 3
                  Perplexity         16        4      80.0%      59.2% 0.1976 +0.3625 4
                  Pooled             94       26     78.3%      71.9% 0.1563 +0.4472 20

Thesis Table 7.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.4 · source-reported

Model-level results for the three GLMRS conditions (relational summaries preloaded, No Clue)

  Condition      Model        Correct Incorrect Accuracy Mean Conf.   Brier       EV SF

  GLMRS-         ChatGPT           17        3    85.0%       62.2%   0.1766   +0.4455   3
  C No Clue
                 Gemini           15        5      75.0%      85.0% 0.2090 +0.4200  5
                 Claude           17        3      85.0%      65.0% 0.1575 +0.4715  3
                 Grok              9       11      45.0%      71.5% 0.3206 −0.0725 11
                 DeepSeek         14        6      70.0%      71.8% 0.1834 +0.3175  6
                 Perplexity        9       11      45.0%      68.5% 0.2863 −0.0450 11
                 Pooled           81       39     67.5%      70.7% 0.2222 +0.2562 39

  GLMRS-         ChatGPT           17        3    85.0%       69.0%   0.1673   +0.4785   3
  D No Clue
                 Gemini           10       10      50.0%      93.0% 0.4218 +0.0150 10
                 Claude           17        3      85.0%      64.0% 0.1615 +0.4635  3
                 Grok              8       12      40.0%      65.9% 0.3030 −0.1180 10
                 DeepSeek         15        5      75.0%      65.5% 0.1710 +0.3600  3
                 Perplexity       12        8      60.0%      72.2% 0.2466 +0.1575  8
                 Pooled           79       41     65.8%      71.6% 0.2452 +0.2261 37

  GLMRS-         ChatGPT           14        6    70.0%       62.0%   0.2200   +0.2530   6
  E No Clue
                 Gemini           13        7      65.0%      84.5% 0.2730 +0.2500  7
                 Claude           17        3      85.0%      62.3% 0.1760 +0.4430  3
                 Grok              8       12      40.0%      66.3% 0.3146 −0.1330 12
                 DeepSeek         14        6      70.0%      72.2% 0.1976 +0.3075  6
                 Perplexity       11        9      55.0%      74.8% 0.2506 +0.1175  9
                 Pooled           77       43     64.2%      70.4% 0.2386 +0.2063 43

Thesis Table 7.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.5 · source-reported

Model accuracy by relational-summary length (No Clue), with preloaded fuller materials for reference

Model         GLMP (mean      GLMRS-C      GLMRS-D        GLMRS-E           E − C
               of runs a,b)     150–300      300–500        500–750

ChatGPT               85.0%       85.0%        85.0%          70.0%        −15.0 pp
Gemini                77.5%       75.0%        50.0%          65.0%        −10.0 pp
Claude                80.0%       85.0%        85.0%          85.0%          0.0 pp
Grok                  60.0%       45.0%        40.0%          40.0%         −5.0 pp
DeepSeek              70.0%       70.0%        75.0%          70.0%          0.0 pp
Perplexity            75.0%       45.0%        60.0%          55.0%        +10.0 pp

Pooled               74.6%       67.5%         65.8%          64.2%        −3.3 pp

Thesis Table 7.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.6 · source-reported

Pooled accuracy: simultaneous loading (first stage) versus preload- ing (second stage)

Comparison                    Simultaneous   Preloaded   Preloaded   Mean of    Mean − si-
                                                 run a       run b   runs a,b   multaneous

Fuller materials, No Clue          87/120       90/120      89/120     74.6%       +2.1 pp
                                  (72.5%)      (75.0%)     (74.2%)
Fuller materials, With Clue        91/120       96/120      94/120     79.2%       +3.3 pp
                                  (75.8%)      (80.0%)     (78.3%)

Thesis Table 7.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.7 · source-reported

Preloaded No Clue comparison by representation and summary length (differences in percentage points)

                                                                         vs GLMP a       vs GLMP b
Condition                                    Correct        Accuracy
                                                                            No Clue         No Clue

GLMP a No Clue (fuller materials)             90/120           75.0%               –              –
GLMP b No Clue (fuller materials,             89/120           74.2%               –              –
rerun)

GLMRS-C (150–300 words)                       81/120           67.5%          −7.5 pp        −6.7 pp
GLMRS-D (300–500 words)                       79/120           65.8%          −9.2 pp        −8.3 pp
GLMRS-E (500–750 words)                       77/120           64.2%         −10.8 pp       −10.0 pp

Thesis Table 7.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.8 · source-reported

Effect of adding the case-specific clue under preloading: model- level accuracy and number of verdicts that changed (out of 20)

                               Run a                                     Run b (rerun)

Model           No Clue      With Clue         Changed         No Clue     With Clue       Changed

ChatGPT            18/20             17/20              1        16/20           18/20            2
Gemini             15/20             15/20              0        16/20           17/20            1
Claude             15/20             17/20              4        17/20           17/20            2
Grok               15/20             20/20              5         9/20            9/20            0
DeepSeek           10/20             10/20              0        18/20           17/20            1
Perplexity         17/20             17/20              2        13/20           16/20            3

Pooled           90/120             96/120             12       89/120         94/120             9

Thesis Table 7.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.9 · source-reported

Run-to-run stability: verdict agreement between run a and the identical rerun b (out of 20 cases)

                         No Clue                               With Clue

Model           Same        Acc. a      Acc. b        Same         Acc. a      Acc. b
               verdict                               verdict

ChatGPT          18/20        90%          80%         17/20          85%        90%
Gemini           15/20        75%          80%         14/20          75%        85%
Claude           18/20        75%          85%         20/20          85%        85%
Grok             12/20        75%          45%          9/20         100%        45%
DeepSeek         12/20        50%          90%         13/20          50%        85%
Perplexity       14/20        85%          65%         17/20          85%        80%

Pooled          89/120      75.0%        74.2%       90/120         80.0%      78.3%
               (74.2%)                              (75.0%)

Thesis Table 7.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.10 · source-reported

Mean stated confidence by model and preloaded six-model condi- tion

             GLMP a          GLMP b
Model
             No Clue GLMP a No Clue GLMP bGLMRS-A    GLMRS-B GLMRS-C GLMRS-D GLMRS-E
                     With Clue      With Clue No Clue No Clue No Clue No Clue No Clue

ChatGPT       67.5%   77.6%   70.7%   79.2%   54.9%   60.4%   62.2%   69.0%   62.0%
Gemini        82.5%   82.5%   87.2%   87.2%   84.5%   84.8%   85.0%   93.0%   84.5%
Claude        67.8%   68.0%   64.8%   65.2%   63.7%   63.8%   65.0%   64.0%   62.3%
Grok          75.8%   93.4%   53.0%   53.0%   65.3%   52.3%   71.5%   65.9%   66.3%
DeepSeek      67.0%   67.0%   78.5%   87.2%   66.8%   71.0%   71.8%   65.5%   72.2%
Perplexity    63.5%   73.9%   59.2%   59.2%   65.0%   66.5%   68.5%   72.2%   74.8%

Pooled        70.7%   77.1%   68.9%   71.9%   66.7%   66.4%   70.7%   71.6%   70.4%

Thesis Table 7.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.11 · source-reported

Sensitivity of the GLMP results (pooled accuracy, and clue effect in percentage points)

Scenario                             a         a         b          b         Clue       Clue
                               No Clue   With Clue No Clue    With Clue    effect a   effect b

As scored (six models)           75.0%      80.0%     74.2%       78.3%       +5.0       +4.2
Excluding Grok (five models,     75.0%      76.0%     80.0%       85.0%       +1.0       +5.0
/100)
Flagged With Clue revisions      75.0%      75.8%     74.2%       75.8%       +0.8       +1.7
reverted∗
      ∗
     Grok run a With Clue replaced by its run-a No Clue verdicts, and Perplexity run b
With Clue replaced by its run-b No Clue verdicts (i.e. the recorded “updated” predictions are
                                      not counted).

Thesis Table 7.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 7.12 · source-reported

Exploratory paired comparisons on the same 120 model–case pairs (discordant pairs and unadjusted exact McNemar p)

Comparison (first vs second)                    First right,   Second right,          Exact p
                                             second wrong        first wrong

GLMP a No Clue vs GLMP b No Clue (rerun)                 16               15             1.000
GLMP a With Clue vs GLMP b With Clue                     16               14             0.856
(rerun)
GLMP a: No Clue vs With Clue                              3                9             0.146
GLMP b: No Clue vs With Clue                              2                7             0.180
GLMP a No Clue vs GLMRS-C                                21               12             0.163
GLMP a No Clue vs GLMRS-D                                23               12             0.090
GLMP a No Clue vs GLMRS-E                                20                7             0.019
GLMP b No Clue vs GLMRS-C                                14                6             0.115
GLMP b No Clue vs GLMRS-D                                16                6             0.052
GLMP b No Clue vs GLMRS-E                                19                7             0.029
GLMRS-C vs GLMRS-D                                       10                8             0.815
GLMRS-D vs GLMRS-E                                       10                8             0.815
GLMRS-C vs GLMRS-E                                       13                9             0.523

These tests treat model–case pairs as independent, although they are clustered by case and by
  model, and no correction is made for the 13 comparisons. The third-stage comparison of
  GLMRS-A with GLMRS-B is reported in Section 8.3. They are reported only to indicate
               which differences are clearly within the range of chance variation.

Thesis Table 7.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.