Key takeaways
- Five preloaded summary lengths yielded 64.2%–67.5%.
- The two shortest lengths each scored 65.8% pooled.
- Individual models followed different observed length curves.
What was tested & why
Stage 3 · Summary length
The plotted points combine experimental stages, and the full-material endpoint is not simply another length of the same summary. Read the chart as observed configurations rather than one controlled dose-response experiment.
Cross-stage points and separately constructed summaries do not establish a causal length effect.
Source: Thesis Chapter 8; Figure 8.1 · source-reported unless otherwise noted.
Results / visual evidence
Observed pooled accuracy by material configuration
Thesis Table 8.4 and Figure 8.1 · points mix experimental stages; this is not a controlled causal curve. Full materials average two runs.
Accuracy across observed configurations
70% reference line. Thesis Table 8.5 / Figure 8.1. The first and last points use different protocols; these are observed configurations, not a causal length curve.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Model-level results for the two third-stage GLMRS conditions (shorter relational summaries preloaded, No Clue)
Condition Model Correct Incorrect Accuracy Mean Conf. Brier EV SF
GLMRS- ChatGPT 15 5 75.0% 54.9% 0.2163 +0.2910 2
A No Clue
Gemini 11 9 55.0% 84.5% 0.3503 +0.0700 9
Claude 15 5 75.0% 63.7% 0.2007 +0.3240 5
Grok 9 11 45.0% 65.3% 0.2976 −0.0625 10
DeepSeek 16 4 80.0% 66.8% 0.1701 +0.4125 4Thesis Table 8.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled results for the two third-stage conditions (120 observations each)
95% Mean Conf.
Condition Correct Acc. Brier EV SF Incorr.
CI∗ conf. −Acc.
GLMRS-A (50–100 w.) 79 65.8% 57.0–73.7 66.7% +0.9 0.2424 +0.2073 37 41
GLMRS-B (100–150 w.) 79 65.8% 57.0–73.7 66.4% +0.6 0.2219 +0.2297 34 41
∗
Wilson interval treating the 120 observations as independent (optimistic, as it ignores
clustering by case and model). SF = Silent Failures under the EV< −0.5 rule. They are fewer
than the incorrect predictions because some incorrect predictions carried confidence of 50% or
less (ChatGPT three and Grok one in GLMRS-A, and Grok six and Perplexity one in
GLMRS-B).Thesis Table 8.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Class-level behaviour in the third-stage conditions (14 Dismissed and 6 Allowed cases per model)
Dismissed Allowed Balanced Predicted
Condition
correct (/84) correct (/36) accuracy Allowed (/120)
GLMRS-A No Clue 62 (73.8%) 17 (47.2%) 60.5% 39
GLMRS-B No Clue 59 (70.2%) 20 (55.6%) 62.9% 45Thesis Table 8.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Pooled summary-length curve (No Clue, 120 observations per point)
Words per
Point Condition Loading Stage Correct Acc. vs 70%
source
0 No Clue (baseline) 0 – Original 81 67.5% −2.5 pp
1 GLMRS-A 50–100 Preloaded Third 79 65.8% −4.2 pp
2 GLMRS-B 100–150 Preloaded Third 79 65.8% −4.2 pp
3 GLMRS-C 150–300 Preloaded Second 81 67.5% −2.5 pp
4 GLMRS-D 300–500 Preloaded Second 79 65.8% −4.2 pp
5 GLMRS-E 500–750 Preloaded Second 77 64.2% −5.8 pp
6 GLMP (runs a, b) Full text Preloaded Second 179/240 74.6% +4.6 pp
The zero point was obtained in the original stage without preloading, and the summary lengths
were run in two different stages. The points therefore describe the observed pattern, not a
single controlled causal curve.Thesis Table 8.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Summary-length curve by model: accuracy (%) over 20 cases
Highest observed
Model 0 A B C D E Full
(length, acc.)
ChatGPT 70.0 75.0 90.0 85.0 85.0 70.0 85.0 100–150 (90.0)
Gemini 60.0 55.0 65.0 75.0 50.0 65.0 77.5 150–300 (75.0)
Claude 70.0 75.0 80.0 85.0 85.0 85.0 80.0 150–300, 300–500, 500–750 (85.0)
Grok 80.0 45.0 50.0 45.0 40.0 40.0 60.0 100–150 (50.0)
DeepSeek 50.0 80.0 70.0 70.0 75.0 70.0 70.0 50–100 (80.0)
Perplexity 75.0 65.0 40.0 45.0 60.0 55.0 75.0 50–100 (65.0)
Pooled 67.5 65.8 65.8 67.5 65.8 64.2 74.6 150–300 (67.5)
Words per source: 0 = original No Clue condition (no added context), A = 50–100, B =
100–150, C = 150–300, D = 300–500 and E = 500–750. Full (GLMP) is the mean of GLMP a
and GLMP b under No Clue. The last column gives the summary length (A–E) at which each
model’s highest observed accuracy occurred. Model-level entries rest on 20 cases (one case = 5
points).Thesis Table 8.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
GLMRS-A (50–100 words) against GLMRS-B (100–150 words): verdict agreement and accuracy change by model
Model Same Acc. A Acc. B Change Right at A Right at B
verdict (pp) only only
ChatGPT 15/20 75.0% 90.0% +15.0 1 4
Gemini 16/20 55.0% 65.0% +10.0 1 3
Claude 17/20 75.0% 80.0% +5.0 1 2
Grok 19/20 45.0% 50.0% +5.0 0 1
DeepSeek 16/20 80.0% 70.0% −10.0 3 1
Perplexity 11/20 65.0% 40.0% −25.0 7 2
Pooled 94/120 65.8% 65.8% 0.0 13 13
(78.3%)Thesis Table 8.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.