Key takeaways
- RQ1–RQ2 address the four-jurisdiction benchmark.
- RQ3–RQ11 address materials, order, length, tools and stability.
- The answer map is in Table 10.14.
What was tested & why
Research questions
The questions below follow the thesis wording and are answered against the observed experiments rather than generalized to every legal task.
Source: Thesis §1.3 · source-reported unless otherwise noted.
The full question set
How do verdict accuracy, confidence reliability and Silent Failures vary across six commercial LLMs when they predict the outcomes of final appellate cases from Nigeria, Australia, the United Kingdom and the United States? Explore evidence
Does adding citation-based case context to the prompt change verdict accuracy and Silent Failures compared with sanitised facts alone, and does the effect vary across models and jurisdictions? Explore evidence
Does adding broad Nigerian legal materials change outcome prediction relative to the original No Clue and With Clue conditions? Explore evidence
When the target case remains under No Clue, does presenting the same general legal materials as connected relational summaries produce different results from presenting the fuller materials? Explore evidence
Do any observed effects remain consistent across the six models? Explore evidence
Does preloading the materials before the prediction task change accuracy relative to supplying them simultaneously with the task? Explore evidence
Are preloaded results stable when the identical experiment is rerun? Explore evidence
Does the permitted length of the relational summaries affect accuracy and reliability? The second stage tested 150–300, 300–500 and 500–750 words per source, and the third stage extended the range downwards to 50–100 and 100–150 words. Explore evidence
How do all conditions compare with the majority-class baseline, and how do results differ across case outcomes and individual cases? Explore evidence
Does each model respond to summary length in the same way, or does each model have its own length curve between no added context and the full materials? Explore evidence
At a fixed summary length, does the tool used to produce the relational summaries change the accuracy, reliability and run-to-run consistency of a fixed prediction model? Explore evidence
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Mapping of the research questions to experiments, results, dis- cussion and conclusion
RQ Topic Part and Results Discussion and
stage(s) conclusion
1 Model, jurisdiction and Part I Ch. 4 Sec. 4.10, 10.6, Ch. 11
confidence (Findings 1 and 2)
2 Citation context Part I, continua- Ch. 4 Sec. 4.10, 10.6, Ch. 11
tion experiments (Findings 1 and 2)
3 Broad materials versus Part II, first and Ch. 6, Sec. 10.2, 10.3 Sec. 10.7, Ch. 11 (Finding 3)
baselines second
4 Relational summaries Part II, first to Ch. 6, 7, 8 Sec. 10.7, Ch. 11 (Finding 4)
versus fuller materials third
5 Consistency across the Part II, first to Sec. 10.2, 10.3 Sec. 10.7, Ch. 11 (Findings 4
six models third and 6)
6 Preloading versus simul- Part II, first and Ch. 7 Sec. 10.7, Ch. 11 (Finding 3)
taneous loading second
7 Run-to-run stability Part II, second Sec. 7.6, 9.4 Sec. 10.8, Ch. 11 (Finding 6)
and fourth
8 Summary length Part II, second Ch. 7, Sec. 8.2 Sec. 10.7, Ch. 11 (Finding 4)
and third
9 Majority class, outcome Part II, first to Sec. 10.4, 10.5 Sec. 10.7, Ch. 11 (Finding 6)
class and cases third
10 Model-specific length Part II, second Sec. 8.2 Sec. 10.7, Ch. 11 (Finding 4)
curves and third
11 Summarisation tool with Part II, fourth Ch. 9 Sec. 10.7, Ch. 11 (Finding 5)
a fixed predictorThesis Table 10.14 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.