CrossLaw/The problem/Why this matters
02 / 26 · The problemThesis §§1.1–1.5; Chapter 2

An answer can look certain and still be wrong.

Why this matters

Key takeaways

  • Appellate-outcome prediction is not a substitute for legal advice.
  • The study addresses gaps in cross-jurisdiction legal AI evaluation.
  • Reasoning-quality review remains in progress.

What was tested & why

Why this matters

Legal practitioners increasingly use AI for research and drafting, yet a convincing answer can hide a wrong verdict or fabricated authority. Most existing legal AI benchmarks focus on the United States, European Union or China, while Nigerian law has received comparatively little systematic evaluation despite growing use. CrossLaw compares verdicts, confidence and the effects of information supplied across four common-law systems, then examines Nigerian cases more closely. The thesis discusses Mata v. Avianca, Mavundla and Handa & Mallick as context, not as benchmark cases.

Source: Thesis §§1.1–1.5; Chapter 2 · source-reported unless otherwise noted.

Thesis §1.4 · intended beneficiaries

Who this work is for

The thesis identifies Nigerian legal professionals as its primary audience, Australian practitioners as secondary beneficiaries, and researchers and benchmark developers as further beneficiaries. The proposal also names law firms, courts, policymakers and AI developers. Their possible uses below are interpretations, not measured real-world impact.

Primary · Nigerian legal practitioners

The need. AI-assisted research is growing, while independent evidence on Nigerian appellate cases and suitable verification practices is limited.

The contribution. The Nigerian results show how the tested models, citation context and different forms of local legal material behaved on the same cases. This can inform questions to ask and what to verify; it is not a current vendor recommendation.

Findings 1–3 and 5–11 · Thesis §§1.1, 1.4, 1.5, 10.12; proposal Impact Statement

Secondary · Australian practitioners, law firms and courts

The need. A plausible citation or confident legal answer can still be wrong, including on Australian legal questions.

The contribution. The Australian comparisons make confident error and jurisdiction-sensitive performance visible, supporting independent checking of authorities and outcomes rather than reliance on a confidence score.

Findings 1–4 and 11 · Thesis §§1.1, 1.4, 1.5; proposal Impact Statement

Wider · researchers, benchmark builders, developers and policymakers

The need. Evaluations concentrated in better-resourced legal systems may miss local variation, calibration problems and unstable reruns.

The contribution. The shared four-jurisdiction design and Nigerian extension offer testable evidence for broader evaluation, including model-by-jurisdiction comparisons, confidence reliability and repeatability. Governance tools are still planned.

Findings 1–4 and 7–11 · Thesis §§1.1, 1.4, 1.5, 10.10–10.12; proposal Impact Statement

Source: thesis §§1.1, 1.4, 10.12 and proposal Impact Statement. No tool-selection recommendation or evidence of changed practice is claimed.

Thesis §1.5 · major findings

Eleven major findings

01

No single model led in every setting

No single model led in every setting. Grok had the highest aggregate accuracy under No Clue (67/80, 83.8%) and Gemini under With Clue (75/80, 93.8%). Accuracy also varied across jurisdictions within individual models. DeepSeek, for example, was correct on 20/20 United Kingdom cases under No Clue but on 9/20 to 10/20 cases in the other jurisdictions.

Thesis §1.5, Tables 4.1–4.3 · source-reported
What this could mean · editorial interpretationNigerian legal practitioners

A score in one country or prompting condition cannot stand in for evidence on Nigerian cases; compare the relevant model and condition before drawing conclusions.

02

Citation context was associated with higher accuracy

Citation context was associated with higher accuracy for five of six models. Pooled accuracy rose from 346/480 (72.1%) to 390/480 (81.3%). Grok was the exception (67/80 to 64/80). The paired change was statistically significant for Gemini (p = 0.0044) and ChatGPT (p = 0.0347).

Thesis §1.5, Table 4.4 · source-reported
What this could mean · editorial interpretationPractitioners across the four jurisdictions

Case-specific citation context changed results, but it did not help every model. A supplied citation is context to check, not independent confirmation.

03

Confident errors were common

Confident errors were common. Of the 960 primary predictions, 222 (23.1%) were Silent Failures, that is incorrect verdicts stated with more than 50% confidence. The count fell from 134 under No Clue to 88 under With Clue. The jurisdictional pattern did not follow a simple resource ordering, since Australia had the highest Silent Failure rate (30.8%).

Thesis §1.5, Tables 4.7–4.10 · source-reported
What this could mean · editorial interpretationLawyers and courts

An authoritative-sounding answer may be wrong. Verify the outcome and cited authority against primary legal sources rather than relying on stated confidence.

04

Accuracy and confidence reliability were separable

Accuracy and confidence reliability were separable. DeepSeek’s 20/20 United Kingdom result coexisted with the highest Expected Calibration Error and Brier score of the six models, and fell to 16/20 on a rerun.

Thesis §1.5, §§4.4, 4.9 · source-reported
What this could mean · editorial interpretationLaw firms evaluating tools

Accuracy alone hides unreliable confidence and rerun instability. Examine calibration and repeated-case results alongside correct-verdict counts.

05

General legal materials helped modestly

General legal materials helped modestly, clues helped most. Adding the fixed 11-source Nigerian corpus raised facts-only (No Clue) accuracy from 67.5% to 72.5% under simultaneous loading and to 75.0% and 74.2% when preloaded, but no material-based condition reached the original With Clue baseline (84.2%).

Thesis §1.5, Chapters 6–7 · source-reported
What this could mean · editorial interpretationNigerian practitioners

Adding general Nigerian legal material had a smaller observed association than case-specific context in these tests; more material alone is not a dependable remedy.

06

Preloading was modestly better

Preloading was modestly better than simultaneous loading (+1.7 to +4.2 points), a consistent direction but within run-to-run variation.

Thesis §1.5, Chapter 7 · source-reported
What this could mean · editorial interpretationTeams designing legal research workflows

Preloading deserves controlled testing, not a blanket recommendation: its modest observed advantage was within run-to-run variation.

07

Model-constructed summaries did not match fuller materials

Model-constructed relational summaries did not match the fuller materials under preloading. Across five preloaded lengths from 50–100 to 500–750 words per source, pooled accuracy was flat (64.2%–67.5%), below the preloaded fuller materials (74.6%) and below the 70.0% always-Dismissed reference.

Thesis §1.5, Chapters 7–8 · source-reported
What this could mean · editorial interpretationTeams preparing Nigerian legal summaries

Check which legal rules a summary retains before treating a shorter, model-written version as equivalent to fuller materials.

08

Each model had its own length curve

Each model had its own length curve. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT (90.0%) and at 150–300 words or longer for Claude, while Grok and Perplexity were below their facts-only baselines at every length. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length.

Thesis §1.5, Chapter 8 · source-reported
What this could mean · editorial interpretationBenchmark designers and developers

Test summary lengths separately for each model; a pooled length result concealed different model-level patterns.

09

The summarisation tool mattered

The summarisation tool mattered when the predictor was fixed. With ChatGPT as the predictor, stored summaries from ChatGPT and Claude gave 90.0% at 100–150 words and 85.0% at 150–300 words, against 80.0% for Gemini Notebook and 75.0% for Microsoft Copilot and Gemini. Every tool gave equal or higher accuracy at the shorter length. Claude’s summaries at 100–150 words gave the highest two-run mean (92.5%, the mean of 90.0% and 95.0%). This is descriptively above ChatGPT’s accuracy with the preloaded full materials in the second stage (85.0%), although that comparison crosses stages and is not controlled.

Thesis §1.5, Chapter 9 · source-reported
What this could mean · editorial interpretationTeams comparing summarisation tools

The tool creating a summary can affect downstream answers even with the predictor held fixed. This small, descriptive comparison needs replication before tool-selection advice.

10

Run-to-run consistency is itself a finding

Run-to-run consistency is itself a finding. Identical reruns of the six-model GLMP experiment agreed on only about three verdicts in four, with swings of up to 55 points for a single model (Grok) and 40 points for another (DeepSeek), whereas the fixed-predictor reruns of the fourth stage agreed on 18–19 of 20 verdicts.

Thesis §1.5, Chapters 7 and 9 · source-reported
What this could mean · editorial interpretationResearchers and procurement teams

Repeat identical trials and report variation; a single run can give a misleading impression of how a system performs.

11

Confidence was not a reliable signal

Confidence was not a reliable signal of correctness, and accuracy in all conditions was markedly lower on Allowed than on Dismissed cases. Two Nigerian cases, NG 009 and NG 020, were predicted incorrectly by all six models under No Clue in the primary benchmark and remained incorrect in every run-a condition of the fourth stage.

Thesis §1.5, Chapters 9–10 · source-reported
What this could mean · editorial interpretationLawyers, courts and evaluators

Check performance on both Allowed and Dismissed outcomes and examine difficult cases; a high aggregate score can conceal asymmetric errors.