Key takeaways
- Appellate-outcome prediction is not a substitute for legal advice.
- The study addresses gaps in cross-jurisdiction legal AI evaluation.
- Reasoning-quality review remains in progress.
What was tested & why
Why this matters
Legal practitioners increasingly use AI for research and drafting, yet a convincing answer can hide a wrong verdict or fabricated authority. Most existing legal AI benchmarks focus on the United States, European Union or China, while Nigerian law has received comparatively little systematic evaluation despite growing use. CrossLaw compares verdicts, confidence and the effects of information supplied across four common-law systems, then examines Nigerian cases more closely. The thesis discusses Mata v. Avianca, Mavundla and Handa & Mallick as context, not as benchmark cases.
Source: Thesis §§1.1–1.5; Chapter 2 · source-reported unless otherwise noted.
Thesis §1.4 · intended beneficiaries
Who this work is for
The thesis identifies Nigerian legal professionals as its primary audience, Australian practitioners as secondary beneficiaries, and researchers and benchmark developers as further beneficiaries. The proposal also names law firms, courts, policymakers and AI developers. Their possible uses below are interpretations, not measured real-world impact.
Primary · Nigerian legal practitioners
The need. AI-assisted research is growing, while independent evidence on Nigerian appellate cases and suitable verification practices is limited.
The contribution. The Nigerian results show how the tested models, citation context and different forms of local legal material behaved on the same cases. This can inform questions to ask and what to verify; it is not a current vendor recommendation.
Findings 1–3 and 5–11 · Thesis §§1.1, 1.4, 1.5, 10.12; proposal Impact StatementSecondary · Australian practitioners, law firms and courts
The need. A plausible citation or confident legal answer can still be wrong, including on Australian legal questions.
The contribution. The Australian comparisons make confident error and jurisdiction-sensitive performance visible, supporting independent checking of authorities and outcomes rather than reliance on a confidence score.
Findings 1–4 and 11 · Thesis §§1.1, 1.4, 1.5; proposal Impact StatementWider · researchers, benchmark builders, developers and policymakers
The need. Evaluations concentrated in better-resourced legal systems may miss local variation, calibration problems and unstable reruns.
The contribution. The shared four-jurisdiction design and Nigerian extension offer testable evidence for broader evaluation, including model-by-jurisdiction comparisons, confidence reliability and repeatability. Governance tools are still planned.
Findings 1–4 and 7–11 · Thesis §§1.1, 1.4, 1.5, 10.10–10.12; proposal Impact StatementSource: thesis §§1.1, 1.4, 10.12 and proposal Impact Statement. No tool-selection recommendation or evidence of changed practice is claimed.
Thesis §1.5 · major findings
Eleven major findings
No single model led in every setting
No single model led in every setting. Grok had the highest aggregate accuracy under No Clue (67/80, 83.8%) and Gemini under With Clue (75/80, 93.8%). Accuracy also varied across jurisdictions within individual models. DeepSeek, for example, was correct on 20/20 United Kingdom cases under No Clue but on 9/20 to 10/20 cases in the other jurisdictions.
Thesis §1.5, Tables 4.1–4.3 · source-reportedA score in one country or prompting condition cannot stand in for evidence on Nigerian cases; compare the relevant model and condition before drawing conclusions.
Citation context was associated with higher accuracy
Citation context was associated with higher accuracy for five of six models. Pooled accuracy rose from 346/480 (72.1%) to 390/480 (81.3%). Grok was the exception (67/80 to 64/80). The paired change was statistically significant for Gemini (p = 0.0044) and ChatGPT (p = 0.0347).
Thesis §1.5, Table 4.4 · source-reportedCase-specific citation context changed results, but it did not help every model. A supplied citation is context to check, not independent confirmation.
Confident errors were common
Confident errors were common. Of the 960 primary predictions, 222 (23.1%) were Silent Failures, that is incorrect verdicts stated with more than 50% confidence. The count fell from 134 under No Clue to 88 under With Clue. The jurisdictional pattern did not follow a simple resource ordering, since Australia had the highest Silent Failure rate (30.8%).
Thesis §1.5, Tables 4.7–4.10 · source-reportedAn authoritative-sounding answer may be wrong. Verify the outcome and cited authority against primary legal sources rather than relying on stated confidence.
Accuracy and confidence reliability were separable
Accuracy and confidence reliability were separable. DeepSeek’s 20/20 United Kingdom result coexisted with the highest Expected Calibration Error and Brier score of the six models, and fell to 16/20 on a rerun.
Thesis §1.5, §§4.4, 4.9 · source-reportedAccuracy alone hides unreliable confidence and rerun instability. Examine calibration and repeated-case results alongside correct-verdict counts.
General legal materials helped modestly
General legal materials helped modestly, clues helped most. Adding the fixed 11-source Nigerian corpus raised facts-only (No Clue) accuracy from 67.5% to 72.5% under simultaneous loading and to 75.0% and 74.2% when preloaded, but no material-based condition reached the original With Clue baseline (84.2%).
Thesis §1.5, Chapters 6–7 · source-reportedAdding general Nigerian legal material had a smaller observed association than case-specific context in these tests; more material alone is not a dependable remedy.
Preloading was modestly better
Preloading was modestly better than simultaneous loading (+1.7 to +4.2 points), a consistent direction but within run-to-run variation.
Thesis §1.5, Chapter 7 · source-reportedPreloading deserves controlled testing, not a blanket recommendation: its modest observed advantage was within run-to-run variation.
Model-constructed summaries did not match fuller materials
Model-constructed relational summaries did not match the fuller materials under preloading. Across five preloaded lengths from 50–100 to 500–750 words per source, pooled accuracy was flat (64.2%–67.5%), below the preloaded fuller materials (74.6%) and below the 70.0% always-Dismissed reference.
Thesis §1.5, Chapters 7–8 · source-reportedCheck which legal rules a summary retains before treating a shorter, model-written version as equivalent to fuller materials.
Each model had its own length curve
Each model had its own length curve. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT (90.0%) and at 150–300 words or longer for Claude, while Grok and Perplexity were below their facts-only baselines at every length. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length.
Thesis §1.5, Chapter 8 · source-reportedTest summary lengths separately for each model; a pooled length result concealed different model-level patterns.
The summarisation tool mattered
The summarisation tool mattered when the predictor was fixed. With ChatGPT as the predictor, stored summaries from ChatGPT and Claude gave 90.0% at 100–150 words and 85.0% at 150–300 words, against 80.0% for Gemini Notebook and 75.0% for Microsoft Copilot and Gemini. Every tool gave equal or higher accuracy at the shorter length. Claude’s summaries at 100–150 words gave the highest two-run mean (92.5%, the mean of 90.0% and 95.0%). This is descriptively above ChatGPT’s accuracy with the preloaded full materials in the second stage (85.0%), although that comparison crosses stages and is not controlled.
Thesis §1.5, Chapter 9 · source-reportedThe tool creating a summary can affect downstream answers even with the predictor held fixed. This small, descriptive comparison needs replication before tool-selection advice.
Run-to-run consistency is itself a finding
Run-to-run consistency is itself a finding. Identical reruns of the six-model GLMP experiment agreed on only about three verdicts in four, with swings of up to 55 points for a single model (Grok) and 40 points for another (DeepSeek), whereas the fixed-predictor reruns of the fourth stage agreed on 18–19 of 20 verdicts.
Thesis §1.5, Chapters 7 and 9 · source-reportedRepeat identical trials and report variation; a single run can give a misleading impression of how a system performs.
Confidence was not a reliable signal
Confidence was not a reliable signal of correctness, and accuracy in all conditions was markedly lower on Allowed than on Dismissed cases. Two Nigerian cases, NG 009 and NG 020, were predicted incorrectly by all six models under No Clue in the primary benchmark and remained incorrect in every run-a condition of the fourth stage.
Thesis §1.5, Chapters 9–10 · source-reportedCheck performance on both Allowed and Dismissed outcomes and examine difficult cases; a high aggregate score can conceal asymmetric errors.