Key takeaways
- RQ1–RQ2: model, jurisdiction, citation and confidence.
- RQ3–RQ10: representation, order, length, baseline and stability.
- RQ11: summarisation-tool differences with a fixed predictor.
What was tested & why
Answers to the 11 research questions
The thesis answers RQ1–RQ11 in §10.9 and maps each question to experiments and results in Table 10.14. Use the linked evidence pages below to inspect each answer.
Source: Thesis §10.9; Table 10.14 · source-reported unless otherwise noted.
Thesis §10.9 · source-reported
Answers and evidence
How do verdict accuracy, confidence reliability and Silent Failures vary across six commercial LLMs when they predict the outcomes of final appellate cases from Nigeria, Australia, the United Kingdom and the United States?
They varied substantially. Grok had the highest aggregate No Clue accuracy (83.8%) and Gemini the highest With Clue accuracy (93.8%), so no model led in both conditions. Models varied across jurisdictions, most clearly DeepSeek (20/20 in the United Kingdom against 9 or 10 of 20 elsewhere under No Clue), and only DeepSeek showed a significant ordered decline from the United Kingdom to Australia to Nigeria. DeepSeek also had the highest ECE and Brier score in both conditions. Of the 960 predictions, 222 (23.1%) were Silent Failures, and the rate was highest for Australia (30.8%), then Nigeria (23.8%), the United States (22.5%) and the United Kingdom (15.4%). Performance therefore did not follow a simple legal-resource hierarchy.
Thesis §10.9 · Table 10.14 · source-reportedDoes adding citation-based case context to the prompt change verdict accuracy and Silent Failures compared with sanitised facts alone, and does the effect vary across models and jurisdictions?
Pooled accuracy rose from 72.1% to 81.3%, and Silent Failures fell from 134 to 88. Five models improved and Grok declined, and the paired change was statistically significant for Gemini and ChatGPT only. The pooled gain was largest for Nigeria (+16.7 points) and smallest for the United States (+3.3 points). The Grok continuation experiment suggested that the presentation of the context also mattered. The effect was therefore model-dependent and did not establish why models changed their answers.
Thesis §10.9 · Table 10.14 · source-reportedDoes adding broad Nigerian legal materials change outcome prediction relative to the original No Clue and With Clue conditions?
Under No Clue, the fuller materials were associated with higher pooled accuracy than facts alone: 72.5% under simultaneous loading and 75.0% and 74.2% when preloaded, against 67.5%. Under With Clue, no material-based condition reached the original With Clue baseline (75.8% to 80.0% against 84.2%). The materials therefore narrowed the gap between the two baselines but did not substitute for the case-specific clue.
Thesis §10.9 · Table 10.14 · source-reportedWhen the target case remains under No Clue, does presenting the same general legal materials as connected relational summaries produce different results from presenting the fuller materials?
Under simultaneous loading they gave the same pooled accuracy (72.5% each). Under preloading, model-constructed summaries at five lengths gave 64.2% to 67.5%, below the preloaded fuller materials (74.6%). The answer therefore depended on the loading protocol and on how the summaries were produced.
Thesis §10.9 · Table 10.14 · source-reportedDo any observed effects remain consistent across the six models?
No. In both stages the same condition was associated with higher accuracy for some models and lower accuracy for others. Grok was below its own factsonly baseline in all nine material-based No Clue conditions and Perplexity in seven of them, whereas ChatGPT, Gemini, Claude and DeepSeek were above their baselines in most conditions.
Thesis §10.9 · Table 10.14 · source-reportedDoes preloading the materials before the prediction task change accuracy relative to supplying them simultaneously with the task?
The preloaded fuller-material runs were 1.7 to 4.2 points above the corresponding simultaneous-loading results. The direction was consistent, but the size was within the observed run-to-run variation, and the comparison crosses stages that were run at different times.
Thesis §10.9 · Table 10.14 · source-reportedAre preloaded results stable when the identical experiment is rerun?
At pooled level, largely yes (75.0% and 74.2% under No Clue, 80.0% and 78.3% under With Clue). At verdict level, agreement was about 75%, with model-level swings of up to 55 points. The fixed-predictor reruns of the fourth stage agreed on 18 or 19 of 20 verdicts. Single-run evaluation may therefore provide an incomplete picture of model behaviour.
Thesis §10.9 · Table 10.14 · source-reportedDoes the permitted length of the relational summaries affect accuracy and reliability? The second stage tested 150–300, 300–500 and 500–750 words per source, and the third stage extended the range downwards to 50–100 and 100–150 words.
No reliable pooled effect was observed. Pooled accuracy was 65.8%, 65.8%, 67.5%, 65.8% and 64.2% from 50–100 to 500–750 words, and no pairwise difference was distinguishable from chance. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length. The confidence–accuracy gap was close to zero at the two shortest lengths and larger at the longer lengths.
Thesis §10.9 · Table 10.14 · source-reportedHow do all conditions compare with the majority-class baseline, and how do results differ across case outcomes and individual cases?
Only the conditions that included the clue and the first preloaded No Clue run exceeded the 70.0% always-Dismissed reference by five points or more. Across the nine preloaded six-model conditions, accuracy was 77.4% on Dismissed cases and 55.2% on Allowed cases. Difficulty was concentrated in a few cases, notably NG 009, NG 006, NG 017, NG 019, NG 020 and NG 011.
Thesis §10.9 · Table 10.14 · source-reportedDoes each model respond to summary length in the same way, or does each model have its own length curve between no added context and the full materials?
No. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT, at 150–300 words for Gemini and at 150– 300 words or longer for Claude. Grok and Perplexity were below their facts-only baselines at every summary length.
Thesis §10.9 · Table 10.14 · source-reportedAt a fixed summary length, does the tool used to produce the relational summaries change the accuracy, reliability and run-to-run consistency of a fixed prediction model?
With ChatGPT as the predictor, accuracy ranged from 75.0% to 90.0% across the five tools at each length, with the highest observed accuracy for ChatGPT’s and Claude’s summaries. The differences arose on five cases and are not statistically separable on 20 cases. Every tool gave equal or higher accuracy at 100–150 words than at 150–300 words. ChatGPT was under-confident in every tool condition, and the reruns of the ChatGPT- and Claude-summary conditions agreed on 18 or 19 of 20 verdicts.
Thesis §10.9 · Table 10.14 · source-reportedSource tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Mapping of the research questions to experiments, results, dis- cussion and conclusion
RQ Topic Part and Results Discussion and
stage(s) conclusion
1 Model, jurisdiction and Part I Ch. 4 Sec. 4.10, 10.6, Ch. 11
confidence (Findings 1 and 2)
2 Citation context Part I, continua- Ch. 4 Sec. 4.10, 10.6, Ch. 11
tion experiments (Findings 1 and 2)
3 Broad materials versus Part II, first and Ch. 6, Sec. 10.2, 10.3 Sec. 10.7, Ch. 11 (Finding 3)
baselines second
4 Relational summaries Part II, first to Ch. 6, 7, 8 Sec. 10.7, Ch. 11 (Finding 4)
versus fuller materials third
5 Consistency across the Part II, first to Sec. 10.2, 10.3 Sec. 10.7, Ch. 11 (Findings 4
six models third and 6)
6 Preloading versus simul- Part II, first and Ch. 7 Sec. 10.7, Ch. 11 (Finding 3)
taneous loading second
7 Run-to-run stability Part II, second Sec. 7.6, 9.4 Sec. 10.8, Ch. 11 (Finding 6)
and fourth
8 Summary length Part II, second Ch. 7, Sec. 8.2 Sec. 10.7, Ch. 11 (Finding 4)
and third
9 Majority class, outcome Part II, first to Sec. 10.4, 10.5 Sec. 10.7, Ch. 11 (Finding 6)
class and cases third
10 Model-specific length Part II, second Sec. 8.2 Sec. 10.7, Ch. 11 (Finding 4)
curves and third
11 Summarisation tool with Part II, fourth Ch. 9 Sec. 10.7, Ch. 11 (Finding 5)
a fixed predictorThesis Table 10.14 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.