CrossLaw/Part II · Nigeria/Stage 3 · Summary length
13 / 26 · Part II · NigeriaThesis Chapter 8; Figure 8.1

More words did not produce a monotonic pooled gain.

Stage 3 · Summary length

Key takeaways

  • Five preloaded summary lengths yielded 64.2%–67.5%.
  • The two shortest lengths each scored 65.8% pooled.
  • Individual models followed different observed length curves.

What was tested & why

Stage 3 · Summary length

The plotted points combine experimental stages, and the full-material endpoint is not simply another length of the same summary. Read the chart as observed configurations rather than one controlled dose-response experiment.

Evidence boundary

Cross-stage points and separately constructed summaries do not establish a causal length effect.

Source: Thesis Chapter 8; Figure 8.1 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Observed pooled accuracy by material configuration

Accuracy— 70% always-Dismissed reference
No added context
67.5%
50–100 words
65.8%
100–150 words
65.8%
150–300 words
67.5%
300–500 words
65.8%
500–750 words
64.2%
Full materials
74.6%

Thesis Table 8.4 and Figure 8.1 · points mix experimental stages; this is not a controlled causal curve. Full materials average two runs.

Figure 8.1 · interactive recreation

Accuracy across observed configurations

70% reference line. Thesis Table 8.5 / Figure 8.1. The first and last points use different protocols; these are observed configurations, not a causal length curve.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 8.1 · source-reported

Model-level results for the two third-stage GLMRS conditions (shorter relational summaries preloaded, No Clue)

  Condition      Model      Correct Incorrect Accuracy Mean Conf.   Brier       EV SF

  GLMRS-         ChatGPT         15        5    75.0%       54.9%   0.2163   +0.2910    2
  A No Clue
                 Gemini          11        9    55.0%       84.5%   0.3503   +0.0700    9
                 Claude          15        5    75.0%       63.7%   0.2007   +0.3240    5
                 Grok             9       11    45.0%       65.3%   0.2976   −0.0625   10
                 DeepSeek        16        4    80.0%       66.8%   0.1701   +0.4125    4

Thesis Table 8.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 8.2 · source-reported

Pooled results for the two third-stage conditions (120 observations each)

                                                 95%    Mean Conf.
       Condition              Correct   Acc.                        Brier          EV SF Incorr.
                                                 CI∗    conf. −Acc.

       GLMRS-A (50–100 w.)          79 65.8% 57.0–73.7 66.7%       +0.9 0.2424 +0.2073     37        41
       GLMRS-B (100–150 w.)         79 65.8% 57.0–73.7 66.4%       +0.6 0.2219 +0.2297     34        41
   ∗
     Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). SF = Silent Failures under the EV< −0.5 rule. They are fewer
than the incorrect predictions because some incorrect predictions carried confidence of 50% or
    less (ChatGPT three and Grok one in GLMRS-A, and Grok six and Perplexity one in
                                         GLMRS-B).

Thesis Table 8.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 8.3 · source-reported

Class-level behaviour in the third-stage conditions (14 Dismissed and 6 Allowed cases per model)

                                Dismissed          Allowed         Balanced       Predicted
          Condition
                               correct (/84)     correct (/36)     accuracy    Allowed (/120)

          GLMRS-A No Clue           62 (73.8%)      17 (47.2%)         60.5%                    39
          GLMRS-B No Clue           59 (70.2%)      20 (55.6%)         62.9%                    45

Thesis Table 8.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 8.4 · source-reported

Pooled summary-length curve (No Clue, 120 observations per point)

                                          Words per
             Point Condition                        Loading         Stage      Correct   Acc. vs 70%
                                          source

              0      No Clue (baseline)   0            –            Original        81   67.5%   −2.5 pp
              1      GLMRS-A              50–100       Preloaded    Third           79   65.8%   −4.2 pp
              2      GLMRS-B              100–150      Preloaded    Third           79   65.8%   −4.2 pp
              3      GLMRS-C              150–300      Preloaded    Second          81   67.5%   −2.5 pp
              4      GLMRS-D              300–500      Preloaded    Second          79   65.8%   −4.2 pp
              5      GLMRS-E              500–750      Preloaded    Second          77   64.2%   −5.8 pp
              6      GLMP (runs a, b)     Full text    Preloaded    Second     179/240   74.6%   +4.6 pp

The zero point was obtained in the original stage without preloading, and the summary lengths
  were run in two different stages. The points therefore describe the observed pattern, not a
                                 single controlled causal curve.

Thesis Table 8.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 8.5 · source-reported

Summary-length curve by model: accuracy (%) over 20 cases

                                                                     Highest observed
Model          0       A       B      C        D       E     Full
                                                                     (length, acc.)

ChatGPT       70.0    75.0    90.0   85.0     85.0    70.0   85.0    100–150 (90.0)
Gemini        60.0    55.0    65.0   75.0     50.0    65.0   77.5    150–300 (75.0)
Claude        70.0    75.0    80.0   85.0     85.0    85.0   80.0    150–300, 300–500, 500–750 (85.0)
Grok          80.0    45.0    50.0   45.0     40.0    40.0   60.0    100–150 (50.0)
DeepSeek      50.0    80.0    70.0   70.0     75.0    70.0   70.0    50–100 (80.0)
Perplexity    75.0    65.0    40.0   45.0     60.0    55.0   75.0    50–100 (65.0)

Pooled        67.5    65.8    65.8   67.5     65.8    64.2   74.6    150–300 (67.5)

   Words per source: 0 = original No Clue condition (no added context), A = 50–100, B =
100–150, C = 150–300, D = 300–500 and E = 500–750. Full (GLMP) is the mean of GLMP a
and GLMP b under No Clue. The last column gives the summary length (A–E) at which each
model’s highest observed accuracy occurred. Model-level entries rest on 20 cases (one case = 5
                                          points).

Thesis Table 8.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 8.6 · source-reported

GLMRS-A (50–100 words) against GLMRS-B (100–150 words): verdict agreement and accuracy change by model

Model            Same     Acc. A      Acc. B      Change    Right at A   Right at B
                verdict                             (pp)          only         only

ChatGPT           15/20    75.0%       90.0%        +15.0            1            4
Gemini            16/20    55.0%       65.0%        +10.0            1            3
Claude            17/20    75.0%       80.0%         +5.0            1            2
Grok              19/20    45.0%       50.0%         +5.0            0            1
DeepSeek          16/20    80.0%       70.0%        −10.0            3            1
Perplexity        11/20    65.0%       40.0%        −25.0            7            2

Pooled          94/120    65.8%        65.8%          0.0          13           13
               (78.3%)

Thesis Table 8.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.