Key takeaways
- Grok led No Clue at 83.8%; Gemini led With Clue at 93.8%.
- Pooled No Clue accuracy was 346/480 (72.1%).
- These results are for specific interfaces and cases, not a vendor ranking.
What was tested & why
Accuracy across models and jurisdictions
Accuracy varied by both model and jurisdiction. DeepSeek scored 20/20 on UK No Clue cases but 9/20–10/20 in the other jurisdictions. This contrast does not establish a general jurisdictional hierarchy.
Source: Thesis §4.2; Tables 4.1–4.3; Figure 4.1 · source-reported unless otherwise noted.
Results / visual evidence
Accuracy by model and prompting condition
Thesis Tables 4.1–4.2 and Figure 4.1 · 80 cases per model per condition. The paired tests are reported in Table 4.4.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
No Clue accuracy: every model in every jurisdiction
| Model | Nigeria | UK | US | Australia | All 80 |
|---|---|---|---|---|---|
| ChatGPT | 70% | 80% | 65% | 60% | 68.8% |
| Gemini | 60% | 75% | 100% | 75% | 77.5% |
| Claude | 70% | 75% | 90% | 60% | 73.8% |
| Grok | 80% | 90% | 90% | 75% | 83.8% |
| DeepSeek | 50% | 100% | 45% | 45% | 60% |
| Perplexity | 75% | 60% | 65% | 75% | 68.8% |
| Pooled | 67.5% | 80% | 75.8% | 65% | 72.1% |
Correct out of 20 per jurisdiction, shown as %. No single model led in every jurisdiction: DeepSeek was 100% on UK cases but 45–50% elsewhere; Gemini was 100% on US cases but 60% on Nigerian cases.
With Clue accuracy: every model in every jurisdiction
| Model | Nigeria | UK | US | Australia | All 80 |
|---|---|---|---|---|---|
| ChatGPT | 100% | 90% | 70% | 70% | 82.5% |
| Gemini | 95% | 90% | 100% | 90% | 93.8% |
| Claude | 85% | 95% | 75% | 85% | 85% |
| Grok | 80% | 80% | 100% | 60% | 80% |
| DeepSeek | 65% | 85% | 60% | 55% | 66.3% |
| Perplexity | 80% | 95% | 70% | 75% | 80% |
| Pooled | 84.2% | 89.2% | 79.2% | 72.5% | 81.3% |
Correct out of 20 per jurisdiction, shown as %. Australia was the weakest pooled jurisdiction in both conditions; Grok fell to 60% on Australian cases with the clue.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Sections 3–6 · Complete score sheets for Nigeria, UK, US and AustraliaReport lines 202–480
SECTION 3: NIGERIA – COMPLETE FINDINGS
3.1 Without Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| NG_001 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_002 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_003 | Appeal Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| NG_004 | Appeal Dismissed | 1 | 1 | 1 | 1 | 0 | 0 |
| NG_005 | Appeal Dismissed | 0 | 1 | 1 | 1 | 0 | 0 |
| NG_006 | Appeal Allowed | 0 | 0 | 0 | 1 | 0 | 1 |
| NG_007 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_008 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_009 | Appeal Dismissed | 0 | 0 | 0 | 0 | 0 | 0 |
| NG_010 | Appeal Dismissed | 1 | 0 | 0 | 1 | 0 | 0 |
| NG_011 | Appeal Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| NG_012 | Appeal Dismissed | 0 | 0 | 0 | 1 | 1 | 1 |
| NG_013 | Appeal Allowed | 0 | 0 | 0 | 1 | 1 | 1 |
| NG_014 | Appeal Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| NG_015 | Appeal Dismissed | 1 | 1 | 1 | 0 | 1 | 1 |
| NG_016 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_017 | Appeal Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| NG_018 | Appeal Dismissed | 1 | 0 | 1 | 1 | 0 | 1 |
| NG_019 | Appeal Allowed | 1 | 0 | 1 | 1 | 1 | 1 |
| NG_020 | Appeal Dismissed | 0 | 0 | 0 | 0 | 0 | 0 |
| TOTAL | 20/20 | 14 | 12 | 14 | 16 | 10 | 15 |
3.2 With Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| NG_001 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_002 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_003 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_004 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_005 | Appeal Dismissed | 1 | 1 | 1 | 0 | 0 | 0 |
| NG_006 | Appeal Allowed | 1 | 0 | 0 | 0 | 0 | 1 |
| NG_007 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_008 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_009 | Appeal Dismissed | 1 | 1 | 0 | 0 | 0 | 0 |
| NG_010 | Appeal Dismissed | 1 | 1 | 1 | 1 | 0 | 0 |
| NG_011 | Appeal Allowed | 1 | 1 | 1 | 0 | 0 | 1 |
| NG_012 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_013 | Appeal Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_014 | Appeal Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_015 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_016 | Appeal Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| NG_017 | Appeal Allowed | 1 | 1 | 0 | 1 | 1 | 1 |
| NG_018 | Appeal Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| NG_019 | Appeal Allowed | 1 | 1 | 1 | 1 | 1 | 0 |
| NG_020 | Appeal Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| TOTAL | 20/20 | 20 | 19 | 17 | 16 | 13 | 16 |
3.3 Nigeria: Summary Statistics
| Model | No Clue Score | No Clue Accuracy | With Clue Score | With Clue Accuracy | Improvement |
|---|---|---|---|---|---|
| ChatGPT | 14/20 | 70.0% | 20/20 | 100.0% | +6 |
| Gemini | 12/20 | 60.0% | 19/20 | 95.0% | +7 |
| Claude | 14/20 | 70.0% | 17/20 | 85.0% | +3 |
| Grok | 16/20 | 80.0% | 16/20 | 80.0% | 0 |
| DeepSeek | 10/20 | 50.0% | 13/20 | 65.0% | +3 |
| Perplexity | 15/20 | 75.0% | 16/20 | 80.0% | +1 |
3.4 Nigeria: Key Findings
- ChatGPT achieves perfect score with clues (20/20) – the only model to do so in Nigeria. This demonstrates exceptional Nigerian legal knowledge recall when citations are provided.
- Gemini shows the largest improvement (+7 cases) with clue, suggesting strong case-recall capability when citations are provided.
- Grok is most reliable without clues (16/20, 80%), but shows zero improvement with clues – indicating that the cases Grok got wrong were genuinely absent from its training data or resistant to recall.
- DeepSeek performs at chance level without clues (10/20, 50%), confirming limited Nigerian legal knowledge in its training data. Critical Risk Signal: DeepSeek's Expected Value turned negative in several Nigerian With-Clue predictions, meaning its confidence-weighted performance would have cost a practitioner more than random guessing. This is detailed in Section 3.5.
- Universal failure on NG_009 (Ethiopian Airlines v Polaris Bank) – all six models missed the statute-bar point, demonstrating systematic procedural blindness. This is the most significant failure in the Nigerian dataset. Even with full citation context, only ChatGPT and Gemini corrected their predictions.
- Universal failure on NG_020 (Atiba v Suberu) – all six models predicted the borrower's appeal would succeed (0/6), missing the parol evidence trap. This is equally significant as NG_009, revealing a systematic weakness in document hierarchy reasoning. With clues, five of six corrected (all except DeepSeek).
- Translation Framework corrected three scores – without it, NG_013 DeepSeek, NG_019 DeepSeek, and NG_019 Perplexity would have been incorrectly scored, affecting final rankings.
Industrial Implication for Nigerian Practice:
ChatGPT with citations is procurement-ready for primary legal research (100% accuracy). DeepSeek must be placed on restricted vendor list pending independent verification protocol due to its 50% unaided accuracy and negative Expected Value pattern (EV approaching -1.0 on multiple cases).
3.5 DeepSeek Risk Signal – Nigeria
DeepSeek's performance in Nigeria exhibits the same negative Expected Value pattern observed in the US dataset. In several With-Clue instances, DeepSeek expressed high confidence on incorrect predictions, generating:
- Expected Value penalties that would have cost a practitioner more than random guessing
- Brier Scores approaching 1.0 on overconfident errors (e.g., NG_005, NG_006, NG_011)
- Zero correction on NG_020 even with full citation context, unlike all other models
Implication for Nigerian practitioners: DeepSeek should not be used for Nigerian legal research without extensive independent verification. Its combination of low accuracy and high confidence on errors creates a "worst of both worlds" risk profile.
Comparative African Governance Evidence:
This risk pattern is not isolated. In South Africa, the Mavundla v MEC [2025] ZAKZPHC 2 case saw a legal team submit seven hallucinated citations – demonstrating that even Africa's most developed legal system is vulnerable. In Kenya, Republic v Public Procurement Administrative Review Board [2026] KEHC 1620 involved counsel relying on AI-generated legal provisions that did not exist. Rwanda offers a constructive counter-model through its National AI Policy 2023, which mandates transparency for AI in public sector decisions. Nigeria currently lacks any binding AI governance law, placing the entire burden of risk mitigation on individual firms.
SECTION 4: UNITED KINGDOM – COMPLETE FINDINGS
4.1 Without Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| UK_001 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_002 | Allowed | 1 | 0 | 1 | 1 | 1 | 1 |
| UK_003 | Allowed | 1 | 1 | 0 | 1 | 1 | 1 |
| UK_004 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_005 | Allowed | 0 | 1 | 1 | 1 | 1 | 0 |
| UK_006 | Dismissed | 1 | 1 | 0 | 0 | 1 | 1 |
| UK_007 | Allowed | 1 | 1 | 1 | 1 | 1 | 0 |
| UK_008 | Dismissed | 1 | 0 | 0 | 1 | 1 | 1 |
| UK_009 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_010 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_011 | Allowed | 0 | 0 | 1 | 1 | 1 | 1 |
| UK_012 | Dismissed | 0 | 1 | 1 | 1 | 1 | 0 |
| UK_013 | Allowed | 0 | 1 | 1 | 0 | 1 | 0 |
| UK_014 | Allowed | 1 | 0 | 0 | 1 | 1 | 0 |
| UK_015 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_016 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| UK_017 | Allowed | 1 | 1 | 0 | 1 | 1 | 1 |
| UK_018 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_019 | Allowed | 1 | 0 | 1 | 1 | 1 | 0 |
| UK_020 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| TOTAL | 12A/8D | 16 | 15 | 15 | 18 | 20 | 12 |
4.2 With Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| UK_001 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_002 | Allowed | 1 | 0 | 1 | 1 | 1 | 1 |
| UK_003 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_004 | Allowed | 0 | 1 | 1 | 0 | 0 | 1 |
| UK_005 | Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| UK_006 | Dismissed | 1 | 0 | 1 | 1 | 1 | 1 |
| UK_007 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_008 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_009 | Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| UK_010 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_011 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_012 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_013 | Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| UK_014 | Allowed | 1 | 1 | 0 | 1 | 0 | 1 |
| UK_015 | Allowed | 0 | 1 | 1 | 1 | 0 | 0 |
| UK_016 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_017 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_018 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_019 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| UK_020 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| TOTAL | 12A/8D | 18 | 18 | 19 | 16 | 17 | 19 |
4.3 UK: Summary Statistics
| Model | No Clue Score | No Clue Accuracy | With Clue Score | With Clue Accuracy | Improvement |
|---|---|---|---|---|---|
| ChatGPT | 16/20 | 80.0% | 18/20 | 90.0% | +2 |
| Gemini | 15/20 | 75.0% | 18/20 | 90.0% | +3 |
| Claude | 15/20 | 75.0% | 19/20 | 95.0% | +4 |
| Grok | 18/20 | 90.0% | 16/20 | 80.0% | -2 |
| DeepSeek | 20/20 | 100.0% | 17/20 | 85.0% | -3 |
| Perplexity | 12/20 | 60.0% | 19/20 | 95.0% | +7 |
4.4 UK: Key Findings
- DeepSeek achieves perfect score without clues (20/20) – unique among all models across all jurisdictions. This demonstrates exceptional UK legal knowledge in its training data, including correct prediction of 2026 judgments.
- Perplexity shows the largest improvement (+7 cases) – from last place (60%) to joint first (95%) with clues, confirming it is a retrieval-first model that performs best when citations are provided.
- Claude achieves joint highest With-Clue score (19/20, 95%) with strong improvement (+4), indicating comprehensive UK case law coverage.
- Grok declines with clues (-2) – the only model to perform worse with citation context, suggesting citation retrieval may introduce noise rather than signal. This is the second jurisdiction where Grok declines.
- DeepSeek regresses with clues (-3) – from perfect score to 17/20, demonstrating that citation clues can override sound doctrinal reasoning with retrieved-but-inaccurate knowledge. DeepSeek expressed 100% confidence on two of its three With-Clue errors (UK_004 and UK_015).
- Seven cases were universally correct across all models in both conditions (UK_001, UK_007, UK_010, UK_016, UK_017, UK_018, UK_020).
Industrial Implication for UK Practice:
DeepSeek's perfect unaided score (20/20) is deceptive – firms should require mandatory citation verification for any DeepSeek-assisted work due to its regression pattern with clues (-3) and overconfidence on errors (100% confidence on UK_004 and UK_015). For UK law with citations, Claude (95%) and Perplexity (95%) are safer choices.
SECTION 5: UNITED STATES – COMPLETE FINDINGS
5.1 Without Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| US_001 | Allowed | 0 | 1 | 1 | 1 | 0 | 0 |
| US_002 | Dismissed | 0 | 1 | 0 | 1 | 0 | 1 |
| US_003 | Dismissed | 0 | 1 | 0 | 1 | 1 | 0 |
| US_004 | Dismissed | 0 | 1 | 1 | 1 | 1 | 0 |
| US_005 | Allowed | 0 | 1 | 1 | 1 | 1 | 0 |
| US_006 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_007 | Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| US_008 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_009 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_010 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_011 | Allowed | 0 | 1 | 1 | 1 | 0 | 0 |
| US_012 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_013 | Dismissed | 1 | 1 | 1 | 0 | 0 | 1 |
| US_014 | Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_015 | Dismissed | 0 | 1 | 1 | 1 | 0 | 0 |
| US_016 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_017 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_018 | Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_019 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| US_020 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| TOTAL | 10A/10D | 13 | 20 | 18 | 18 | 9 | 13 |
5.2 With Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| US_001 | Allowed | 0 | 1 | 0 | 1 | 0 | 0 |
| US_002 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_003 | Dismissed | 1 | 1 | 0 | 1 | 1 | 0 |
| US_004 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| US_005 | Allowed | 0 | 1 | 0 | 1 | 1 | 1 |
| US_006 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_007 | Allowed | 1 | 1 | 1 | 1 | 1 | 0 |
| US_008 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_009 | Allowed | 0 | 1 | 1 | 1 | 1 | 1 |
| US_010 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_011 | Allowed | 0 | 1 | 1 | 1 | 0 | 1 |
| US_012 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_013 | Dismissed | 0 | 1 | 0 | 1 | 0 | 0 |
| US_014 | Allowed | 0 | 1 | 0 | 1 | 0 | 1 |
| US_015 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_016 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| US_017 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_018 | Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| US_019 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| US_020 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| TOTAL | 10A/10D | 14 | 20 | 15 | 20 | 12 | 14 |
5.3 US: Summary Statistics
| Model | No Clue Score | No Clue Accuracy | With Clue Score | With Clue Accuracy | Improvement |
|---|---|---|---|---|---|
| ChatGPT | 13/20 | 65.0% | 14/20 | 70.0% | +1 |
| Gemini | 20/20 | 100.0% | 20/20 | 100.0% | 0 |
| Claude | 18/20 | 90.0% | 15/20 | 75.0% | -3 |
| Grok | 18/20 | 90.0% | 20/20 | 100.0% | +2 |
| DeepSeek | 9/20 | 45.0% | 12/20 | 60.0% | +3 |
| Perplexity | 13/20 | 65.0% | 14/20 | 70.0% | +1 |
5.4 US: Key Findings
- Gemini achieves perfect score in both conditions (20/20) – the only model to achieve perfect accuracy in both No-Clue and With-Clue in any jurisdiction. Its confidence calibration is exceptional (Avg Brier 0.0025 in No-Clue).
- Grok achieves perfect score with clues (20/20), improving from 18/20 without clues – demonstrating strong US case law retrieval capability.
- Claude regresses with clues (-3) – the only model to decline, suggesting citation-induced knowledge interference. This is a notable negative finding warranting further investigation.
- DeepSeek performs near-random without clues (9/20, 45%) – below chance level on a balanced 10/10 split, confirming its training data is not US-focused.
- US_013 (Medical Marijuana v Horn) was most difficult – only Gemini and Grok correct in With-Clue condition, demonstrating the frontier of recent (2025) jurisprudence. Even with citation, 4 of 6 models failed.
- Cases where all models agreed – US_006 (Cuozzo) and US_009 (Becerra) were correctly predicted by all six models in both conditions.
Industrial Implication for US Practice:
Gemini is the only model that can be deployed without conditional prompting (100% accuracy both conditions). Claude requires citation-specific testing protocols due to its regression with clues (-3). For firms handling US-Nigeria cross-border work, ChatGPT with citations (100% on Nigeria, 70% on US) requires jurisdictional switching protocols.
SECTION 6: AUSTRALIA – COMPLETE FINDINGS
6.1 Without Clue – Complete Score Summary
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| AU_001 | Allowed | 1 | 1 | 1 | 1 | 1 | 0 |
| AU_002 | Dismissed | 0 | 0 | 0 | 1 | 1 | 1 |
| AU_003 | Allowed | 0 | 0 | 0 | 1 | 0 | 1 |
| AU_004 | Dismissed | 1 | 0 | 1 | 1 | 0 | 1 |
| AU_005 | Dismissed | 1 | 1 | 0 | 1 | 0 | 1 |
| AU_006 | Dismissed | 0 | 0 | 0 | 1 | 0 | 1 |
| AU_007 | Allowed | 0 | 1 | 1 | 1 | 0 | 1 |
| AU_008 | Dismissed | 1 | 1 | 1 | 1 | 1 | 0 |
| AU_009 | Allowed | 0 | 1 | 1 | 1 | 0 | 1 |
| AU_010 | Dismissed | 1 | 1 | 1 | 1 | 0 | 1 |
| AU_011 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_012 | Dismissed | 0 | 1 | 0 | 0 | 0 | 0 |
| AU_013 | Allowed | 0 | 0 | 0 | 0 | 0 | 0 |
| AU_014 | Dismissed | 0 | 1 | 0 | 0 | 0 | 0 |
| AU_015 | Allowed | 1 | 1 | 1 | 1 | 0 | 1 |
| AU_016 | Dismissed | 1 | 1 | 1 | 0 | 1 | 1 |
| AU_017 | Allowed | 1 | 1 | 0 | 1 | 1 | 1 |
| AU_018 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_019 | Allowed | 1 | 1 | 1 | 0 | 1 | 1 |
| AU_020 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| TOTAL | — | 12 | 15 | 12 | 15 | 9 | 15 |
6.2 With Clue – Complete Score Summary (CORRECTED)
| Case ID | Ground Truth | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| AU_001 | Allowed | 0 | 1 | 1 | 1 | 0 | 0 |
| AU_002 | Dismissed | 0 | 1 | 0 | 1 | 1 | 1 |
| AU_003 | Allowed | 1 | 1 | 1 | 1 | 0 | 0 |
| AU_004 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_005 | Dismissed | 1 | 0 | 1 | 0 | 0 | 1 |
| AU_006 | Dismissed | 0 | 1 | 1 | 0 | 0 | 1 |
| AU_007 | Allowed | 0 | 1 | 1 | 0 | 1 | 1 |
| AU_008 | Dismissed | 1 | 1 | 0 | 1 | 1 | 1 |
| AU_009 | Allowed | 1 | 1 | 1 | 0 | 0 | 1 |
| AU_010 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_011 | Allowed | 1 | 0 | 1 | 0 | 1 | 1 |
| AU_012 | Dismissed | 0 | 1 | 1 | 1 | 1 | 1 |
| AU_013 | Allowed | 1 | 1 | 0 | 1 | 0 | 0 |
| AU_014 | Dismissed | 0 | 1 | 1 | 0 | 0 | 1 |
| AU_015 | Allowed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_016 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| AU_017 | Allowed | 1 | 1 | 1 | 1 | 0 | 0 |
| AU_018 | Dismissed | 1 | 1 | 1 | 0 | 1 | 1 |
| AU_019 | Allowed | 1 | 1 | 1 | 0 | 0 | 0 |
| AU_020 | Dismissed | 1 | 1 | 1 | 1 | 1 | 1 |
| TOTAL | — | 14 | 18 | 17 | 12 | 11 | 15 |
6.3 Australia: Summary Statistics (CORRECTED)
| Model | No Clue Score | No Clue Accuracy | With Clue Score | With Clue Accuracy | Improvement |
|---|---|---|---|---|---|
| ChatGPT | 12/20 | 60.0% | 14/20 | 70.0% | +2 |
| Gemini | 15/20 | 75.0% | 18/20 | 90.0% | +3 |
| Claude | 12/20 | 60.0% | 17/20 | 85.0% | +5 |
| Grok | 15/20 | 75.0% | 12/20 | 60.0% | -3 |
| DeepSeek | 9/20 | 45.0% | 11/20 | 55.0% | +2 |
| Perplexity | 15/20 | 75.0% | 15/20 | 75.0% | 0 |
6.4 Australia: Key Findings (UPDATED)
- Claude shows largest improvement (+5 cases) – from 12/20 to 17/20 with clues, demonstrating strong responsiveness to Australian citation context.
- Gemini achieves highest With-Clue score (18/20, 90%) – most reliable for Australian law with citation context.
- Grok declines with clues (-3) – now declines in three jurisdictions (UK -2, Australia -3, Nigeria 0), confirming a clear pattern of citation-induced performance degradation. This is the most significant model-specific negative finding.
- AU_013 (Ecosse Property Holdings) produced universal failure (0/6) without clues – all models applied literal textual construction of "payable by the tenant" instead of purposive "de facto sale" interpretation. With clues, three models corrected (ChatGPT, Gemini, Grok).
- AU_012 (Fischer v Nemeske) improved dramatically – from 1/6 (Gemini only) to 5/6 with clues, showing citation retrieval corrects trust law errors.
- Three cases produced perfect 6/6 scores in No Clue condition: AU_011 (Simic), AU_018 (Westpac), and AU_020 (Productivity Partners). These were cases with well-established doctrinal principles well-represented in training data.
- Partial verdict adjudications required – seven instances across both conditions, with Hard Boundary Rule applied consistently.
- Correction Note: The AU_018 Grok (With Clue) score was corrected from 1 → 0, reducing Grok's With-Clue total from 13 to 12 and its accuracy from 65% to 60%. This correction has been applied consistently throughout the report.
Industrial Implication for Australian Practice:
Grok should be deployed with a mandatory "Partial Clue" protocol (year + court level only, no full citations) to optimise its 90% accuracy potential (discovered in Experiment 1, Section EXP-1). Without this protocol, Grok's accuracy falls to 65% with full citations. Gemini remains the most reliable for Australian law with citations (90%).
Sections 8–9 · Model specialisation by jurisdiction and Commonwealth lineageReport lines 530–577
SECTION 8: JURISDICTIONAL BIAS ANALYSIS
8.1 Model-Specialization Patterns (No Clue)
| Model | Best Jurisdiction | Worst Jurisdiction | Gap | Interpretation |
|---|---|---|---|---|
| ChatGPT | UK (80%) | Australia (60%) | 20% | Balanced |
| Gemini | US (100%) | Nigeria (60%) | 40% | Strong US bias |
| Claude | US (90%) | Australia (60%) | 30% | Moderate US bias |
| Grok | UK/US (90%) | Australia (75%) | 15% | Balanced |
| DeepSeek | UK (100%) | US/Australia (45%) | 55% | Extreme UK bias |
| Perplexity | Nigeria/Australia (75%) | UK (60%) | 15% | Slight Commonwealth preference |
8.2 Key Findings on Jurisdictional Bias
- Gemini shows strongest US bias – 40% gap between US (100%) and Nigeria (60%). This suggests US-centric training data with limited Nigerian legal content.
- DeepSeek shows extreme UK bias – 55% gap between UK (100%) and US/Australia (45%). This is the most dramatic jurisdictional bias observed.
- ChatGPT is most jurisdiction-agnostic – range of only 20% across all four jurisdictions, making it the most balanced model.
- Perplexity performs best on Nigeria and Australia – the opposite of expected bias pattern, suggesting different training data composition.
- Grok shows balanced performance – only 15% range, making it the second most jurisdiction-independent model after ChatGPT.
- The "US/UK vs Nigeria/Australia" hypothesis is rejected – patterns are model-specific, not uniform. Some models (DeepSeek) show extreme UK bias, others (Perplexity) show reverse bias.
8.3 Bias Mitigation Procurement Protocol:
For IT managers selecting LLM vendors for Nigerian legal practice:
- Vendor Disclosure Requirement – Require jurisdiction-specific accuracy claims to be verified against independent benchmarks (this study). Vendors must disclose training data composition by jurisdiction.
- Bias Testing Mandate – Include jurisdiction sensitivity testing (Chi-square as per Section 2 Part 2) in vendor evaluation. A statistically significant gap (>0.05 Chi-square p-value) between US and Nigerian performance triggers enhanced verification requirements.
- Contractual Clause – Material performance degradation (>15% gap between advertised and actual Nigerian accuracy) constitutes breach, allowing firm to terminate or demand remediation.
- Ongoing Monitoring – Quarterly re-benchmarking on held-out Nigerian cases from the master dataset (20 cases available upon request). Maintain restricted vendor list for models exhibiting negative Expected Value.
- Example Application: DeepSeek's 40-55% jurisdictional gaps and negative EV trigger automatic placement on restricted list, requiring case-by-case partner approval for any use.
SECTION 9: COMMONWEALTH LINEAGE ANALYSIS
9.1 UK → Australia → Nigeria Trend (No Clue)
| Model | UK | Australia | Nigeria | Trend | Significance |
|---|---|---|---|---|---|
| ChatGPT | 80% | 60% | 70% | Mixed | No clear trend |
| Gemini | 75% | 75% | 60% | ↓ Nigeria only | UK→AU stable, NG drop |
| Claude | 75% | 60% | 70% | Mixed | No clear trend |
| Grok | 90% | 75% | 80% | Mixed | No clear trend |
| DeepSeek | 100% | 45% | 50% | ↓↓ Sharp drop | Strong lineage effect |
| Perplexity | 60% | 75% | 75% | ↑ Increasing | Reverse lineage |
9.2 Key Findings on Commonwealth Lineage
- Only DeepSeek shows the expected Commonwealth degradation pattern – perfect UK knowledge collapses on Australia and Nigeria, suggesting training data hierarchy (UK → Commonwealth → rest).
- Perplexity shows reverse trend – performs best on Nigeria and Australia, worst on UK, indicating different training data composition focused on developing common law jurisdictions.
- Gemini shows stable performance across UK and Australia (both 75%), with drop only on Nigeria, suggesting specific gap in Nigerian legal knowledge rather than general Commonwealth degradation.
- The Commonwealth lineage hypothesis is model-dependent – not a universal phenomenon; some models (ChatGPT, Claude, Grok) show no clear lineage effect.
- This finding has implications for legal practitioners – models cannot be assumed to transfer knowledge across Commonwealth jurisdictions uniformly.
Section 10 · Performance by legal domainReport lines 578–632
SECTION 10: ANALYSIS BY LEGAL DOMAIN
10.1 Nigeria – Performance by Legal Domain
| Domain | Cases | Avg Accuracy (NC) | Risk Level | Verification Requirement | Notable Findings |
|---|---|---|---|---|---|
| Contract Law (Labour) | NG_001, 005, 007 | 77.8% | MEDIUM | Verify reasoning (Type II risk) | NG_001 "correct guessing" |
| Contract Law (Land/Mortgage) | NG_004, 012, 013, 020 | 58.3% | HIGH | Mandatory manual verification of document hierarchy | NG_020 universal failure |
| Tort Law | NG_002, 008, 014-015, 017 | 72.0% | MEDIUM | Standard verification of case law | Strong overall performance |
| Commercial (IP/Banking/Admiralty) | NG_003, 009-011, 018-019 | 66.7% | HIGH | Mandatory procedural check for limitation periods | NG_009 universal failure |
| Criminal Law | NG_016 | 100% | LOW | Routine verification acceptable | All models correct |
Domain Insight: Contract law involving land/mortgage proved most challenging (58.3%), while tort law showed stronger performance (72.0%). Procedural cases (NG_009) and document hierarchy cases (NG_020) exposed systematic weaknesses.
10.2 UK – Performance by Legal Domain
| Domain | Cases | Avg Accuracy (NC) | Risk Level | Verification Requirement | Notable Findings |
|---|---|---|---|---|---|
| IP/Patent | UK_018 | 100% | LOW | Routine verification | Universal correctness |
| Civil Procedure | UK_009 | 100% | LOW | Routine verification | Universal correctness |
| Contract Law | UK_002, 017, 020 | 88.9% | LOW | Routine verification | Strong performance |
| Banking/Commercial | UK_003, 012, 016 | 88.9% | LOW | Routine verification | Strong performance |
| Employment | UK_015 | 83.3% | MEDIUM | Standard verification | Contested in With-Clue |
| Tort Law | UK_001, 004-006, 008, 010-011 | 76.2% | MEDIUM | Verify recent (2026) case logic | UK_004 (2026) contested |
| Property/Equity | UK_013, 014 | 58.3% | HIGH | Mandatory manual verification | UK_013 difficult |
Domain Insight: Property/equity cases proved most challenging (58.3%), while contract, banking, and IP showed strong performance. Scottish private law (UK_006) exposed knowledge gaps.
10.3 US – Performance by Legal Domain
| Domain | Cases | Avg Accuracy (NC) | Risk Level | Verification Requirement | Notable Findings |
|---|---|---|---|---|---|
| Administrative Law | US_009 | 100% | LOW | Routine verification | Universal correctness |
| Employment/Arbitration | US_003, 014, 018, 020 | 87.5% | LOW | Routine verification | Strong performance |
| Banking/Preemption | US_007 | 83.3% | MEDIUM | Standard verification | Moderate difficulty |
| Intellectual Property | US_001, 005-006, 008, 015-017 | 79.8% | MEDIUM | Verify against modern statutes | Well-documented doctrine |
| Bankruptcy | US_010, 012 | 75.0% | MEDIUM | Standard verification | Mixed performance |
| Civil Procedure | US_002, 004, 011 | 66.7% | HIGH | Manual check of jurisdiction rules | US_011 (Mallory) difficult |
| RICO Civil Law | US_013 | 50.0% | HIGH | Mandatory manual verification | Most difficult (2025) |
Domain Insight: RICO (US_013) proved most challenging (50%), reflecting recency effect (2025). Intellectual Property cases showed strong performance (79.8%), suggesting well-documented doctrine in training data.
10.4 Australia – Performance by Legal Domain (UPDATED)
| Domain | Cases | Avg Accuracy (NC) | Risk Level | Verification Requirement | Notable Findings |
|---|---|---|---|---|---|
| Commercial Law | AU_018-020 | 88.9% | LOW | Routine verification | Strong performance |
| Public International Law | AU_008 | 83.3% | MEDIUM | Standard verification | Strong performance |
| IP Law | AU_004, 017 | 75.0% | MEDIUM | Standard verification | Strong performance |
| Tort Law | AU_009, 015 | 66.7% | HIGH | Manual check of duty of care | Moderate performance |
| Contract Law | AU_001, 005, 010-011, 013-014 | 63.9% | HIGH | Manual check of purposive intent | AU_013 universal failure |
| Equity | AU_002, 012, 016 | 58.3% | HIGH | Mandatory manual verification | AU_012 difficult |
| Civil Procedure | AU_006, 007 | 58.3% | HIGH | Manual check of court rules | AU_006 difficult |
| Property Law | AU_003 | 50.0% | HIGH | Mandatory manual verification | Mixed performance |
Domain Insight: Equity cases proved most challenging (58.3%), while commercial law showed strongest performance (88.9%). Contract interpretation cases (AU_013) revealed systematic weaknesses in purposive construction.
Domain Risk Map Summary:
For Nigerian practitioners, Criminal Law (Green/Low Risk) can be safely delegated to AI with routine verification. Tort and non-land contract (Yellow/Medium Risk) require standard verification including reasoning checks. Land/mortgage contract, commercial banking, and IP (Orange/High Risk) require mandatory manual verification. Procedural/limitation and document hierarchy (Red/Critical Risk) should never be delegated to AI – manual verification is mandatory regardless of model confidence.
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Verdict accuracy by model and jurisdiction under No Clue (correct out of 20 per jurisdiction)
Model Nigeria UK US Australia Total Accuracy ChatGPT 14 16 13 12 55/80 68.8% Gemini 12 15 20 15 62/80 77.5% Claude 14 15 18 12 59/80 73.8% Grok 16 18 18 15 67/80 83.8% DeepSeek 10 20 9 9 48/80 60.0% Perplexity 15 12 13 15 55/80 68.8% Pooled 81/120 96/120 91/120 78/120 346/480 72.1%
Thesis Table 4.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Verdict accuracy by model and jurisdiction under With Clue (cor- rect out of 20 per jurisdiction)
Model Nigeria UK US Australia Total Accuracy
ChatGPT 20 18 14 14 66/80 82.5%
Gemini 19 18 20 18 75/80 93.8%
Claude 17 19 15 17 68/80 85.0%
Grok 16 16 20 12 64/80 80.0%
DeepSeek 13 17 12 11 53/80 66.3%
Perplexity 16 19 14 15 64/80 80.0%
Pooled 101/120 107/120 95/120 87/120 390/480 81.3%
The Nigeria columns are the original No Clue (81/120, 67.5%) and With Clue (101/120,
84.2%) baselines of the Nigerian extension.Thesis Table 4.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Commonwealth-lineage comparison under No Clue (Jonckheere– Terpstra test of the ordering UK, Australia, Nigeria)
Model UK Australia Nigeria Observed pattern p
ChatGPT 80.0% 60.0% 70.0% Mixed 0.5874
Gemini 75.0% 75.0% 60.0% Lower Nigerian result 0.4157
Claude 75.0% 60.0% 70.0% Mixed 0.7861
Grok 90.0% 75.0% 80.0% Mixed 0.5874
DeepSeek 100.0% 45.0% 50.0% Marked decline 0.0067
Perplexity 60.0% 75.0% 75.0% Reverse pattern 0.4157Thesis Table 4.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.