Key takeaways
- Performance depended on model, jurisdiction and information supplied.
- Confident errors were common.
- The findings warrant controlled replication and human verification, not universal recommendations.
What was tested & why
Implications & conclusion
The concluding statement is reproduced below from thesis §11.4. It qualifies all result comparisons and does not establish that summaries or preloading always outperform fuller legal materials.
Source: Thesis §§10.12, 11.1–11.4 · source-reported unless otherwise noted.
Thesis §11.3
Six principal findings
Performance depended on the model, the jurisdiction and the prompting condition. No model led in both conditions, and citation context was associated with higher accuracy for five of six models and in all four jurisdictions at pooled level, with the largest pooled gain on Nigerian cases.
Confident error was common and did not follow a simple resource hierarchy. Almost every incorrect primary prediction was stated with more than 50% confidence. Silent Failures were fewer with citation context but remained frequent, and Australia had the highest rate.
General legal materials were associated with modest gains but did not replace case-specific context. Under No Clue, the fuller materials were associated with accuracy 5.0 to 7.5 points above facts alone. No material-based condition reached the original With Clue baseline (84.2%). When the materials were present, the clue was associated with only 3.3 to 5.0 additional points, against 16.7 points in the original pair.
Model-constructed summaries did not reproduce the fuller-material result under preloading, and summary length had no reliable pooled effect. Across five preloaded lengths, pooled accuracy stayed between 64.2% and 67.5%, below the preloaded fuller materials (74.6%). Each model responded to length differently, so any recommendation about summary length would need to be model-specific.
With a fixed predictor, the summarisation tool was associated with differences in accuracy. Accuracy ranged from 75.0% to 90.0% across tools, and shorter summaries gave equal or higher accuracy for every tool. The highest two-run mean (92.5%, Claude’s summaries at 100–150 words) was descriptively above ChatGPT’s second-stage accuracy with the full materials (85.0%). This cross-stage comparison is not controlled, and the tool differences rested on five cases and are not statistically separable on this sample.
Consistency and confidence need to be measured, not assumed. Identical six-model reruns agreed on about three verdicts in four, whereas the fixed-predictor reruns agreed on 18 or 19 of 20. Stated confidence did not track which conditions were more accurate, and accuracy was markedly lower on Allowed than on Dismissed cases. Single-run evaluation may therefore provide an incomplete picture, and repeated runs and human verification should be considered in legal AI evaluation.
For the appellate cases and commercial language models tested, outcome-prediction performance depended on the model, the jurisdiction and the information supplied in the prompt, and confident errors were common. For the Nigerian cases, the representation and order of presentation of general Nigerian legal materials, and the tool used to summarise them, were associated with differences in performance. These differences were small relative to the effect of case-specific citation context, to the majority-class reference, to model-level heterogeneity and to run-to-run variation in the six-model experiments. Preloading the fuller materials was modestly associated with higher accuracy. Model-constructed relational summaries did not reproduce the first-stage parity with fuller materials at any of five lengths, and each model responded to summary length differently. With a fixed predictor, the summarisation tool was associated with differences of up to 15 points, and short stored summaries from ChatGPT and Claude were associated with accuracy at or above the predictor’s full-material result. Stated confidence was not a reliable indicator of correctness. The findings support further controlled study, with repeated runs and audited representations, but they do not establish that summaries or preloading are universally more effective than supplying fuller legal materials.