CrossLaw/Part I · Benchmark/Accuracy across models and jurisdictions
05 / 26 · Part I · BenchmarkThesis §4.2; Tables 4.1–4.3; Figure 4.1

The leading model changed with the prompting condition.

Accuracy across models and jurisdictions

Key takeaways

  • Grok led No Clue at 83.8%; Gemini led With Clue at 93.8%.
  • Pooled No Clue accuracy was 346/480 (72.1%).
  • These results are for specific interfaces and cases, not a vendor ranking.

What was tested & why

Accuracy across models and jurisdictions

Accuracy varied by both model and jurisdiction. DeepSeek scored 20/20 on UK No Clue cases but 9/20–10/20 in the other jurisdictions. This contrast does not establish a general jurisdictional hierarchy.

Source: Thesis §4.2; Tables 4.1–4.3; Figure 4.1 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Accuracy by model and prompting condition

No ClueWith Clue
ChatGPT
68.8%
82.5%
Gemini
77.5%
93.8%
Claude
73.8%
85.0%
Grok
83.8%
80.0%
DeepSeek
60.0%
66.3%
Perplexity
68.8%
80.0%

Thesis Tables 4.1–4.2 and Figure 4.1 · 80 cases per model per condition. The paired tests are reported in Table 4.4.

Model-level comparison

All six models, side by side

Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.

Thesis Table 4.1 · model × jurisdiction · source-reported

No Clue accuracy: every model in every jurisdiction

ModelNigeriaUKUSAustraliaAll 80
ChatGPT70%80%65%60%68.8%
Gemini60%75%100%75%77.5%
Claude70%75%90%60%73.8%
Grok80%90%90%75%83.8%
DeepSeek50%100%45%45%60%
Perplexity75%60%65%75%68.8%
Pooled67.5%80%75.8%65%72.1%

Correct out of 20 per jurisdiction, shown as %. No single model led in every jurisdiction: DeepSeek was 100% on UK cases but 45–50% elsewhere; Gemini was 100% on US cases but 60% on Nigerian cases.

Thesis Table 4.2 · model × jurisdiction · source-reported

With Clue accuracy: every model in every jurisdiction

ModelNigeriaUKUSAustraliaAll 80
ChatGPT100%90%70%70%82.5%
Gemini95%90%100%90%93.8%
Claude85%95%75%85%85%
Grok80%80%100%60%80%
DeepSeek65%85%60%55%66.3%
Perplexity80%95%70%75%80%
Pooled84.2%89.2%79.2%72.5%81.3%

Correct out of 20 per jurisdiction, shown as %. Australia was the weakest pooled jurisdiction in both conditions; Grok fell to 60% on Australian cases with the clue.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Sections 3–6 · Complete score sheets for Nigeria, UK, US and AustraliaReport lines 202–480

SECTION 3: NIGERIA – COMPLETE FINDINGS

3.1 Without Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
NG_001Appeal Dismissed111111
NG_002Appeal Dismissed111111
NG_003Appeal Dismissed111101
NG_004Appeal Dismissed111100
NG_005Appeal Dismissed011100
NG_006Appeal Allowed000101
NG_007Appeal Dismissed111111
NG_008Appeal Dismissed111111
NG_009Appeal Dismissed000000
NG_010Appeal Dismissed100100
NG_011Appeal Allowed111101
NG_012Appeal Dismissed000111
NG_013Appeal Allowed000111
NG_014Appeal Allowed111101
NG_015Appeal Dismissed111011
NG_016Appeal Dismissed111111
NG_017Appeal Allowed111011
NG_018Appeal Dismissed101101
NG_019Appeal Allowed101111
NG_020Appeal Dismissed000000
TOTAL20/20141214161015

3.2 With Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
NG_001Appeal Dismissed111111
NG_002Appeal Dismissed111111
NG_003Appeal Dismissed111111
NG_004Appeal Dismissed111111
NG_005Appeal Dismissed111000
NG_006Appeal Allowed100001
NG_007Appeal Dismissed111111
NG_008Appeal Dismissed111111
NG_009Appeal Dismissed110000
NG_010Appeal Dismissed111100
NG_011Appeal Allowed111001
NG_012Appeal Dismissed111111
NG_013Appeal Allowed111111
NG_014Appeal Allowed111111
NG_015Appeal Dismissed111111
NG_016Appeal Dismissed111111
NG_017Appeal Allowed110111
NG_018Appeal Dismissed111101
NG_019Appeal Allowed111110
NG_020Appeal Dismissed111101
TOTAL20/20201917161316

3.3 Nigeria: Summary Statistics

ModelNo Clue ScoreNo Clue AccuracyWith Clue ScoreWith Clue AccuracyImprovement
ChatGPT14/2070.0%20/20100.0%+6
Gemini12/2060.0%19/2095.0%+7
Claude14/2070.0%17/2085.0%+3
Grok16/2080.0%16/2080.0%0
DeepSeek10/2050.0%13/2065.0%+3
Perplexity15/2075.0%16/2080.0%+1

3.4 Nigeria: Key Findings

  • ChatGPT achieves perfect score with clues (20/20) – the only model to do so in Nigeria. This demonstrates exceptional Nigerian legal knowledge recall when citations are provided.
  • Gemini shows the largest improvement (+7 cases) with clue, suggesting strong case-recall capability when citations are provided.
  • Grok is most reliable without clues (16/20, 80%), but shows zero improvement with clues – indicating that the cases Grok got wrong were genuinely absent from its training data or resistant to recall.
  • DeepSeek performs at chance level without clues (10/20, 50%), confirming limited Nigerian legal knowledge in its training data. Critical Risk Signal: DeepSeek's Expected Value turned negative in several Nigerian With-Clue predictions, meaning its confidence-weighted performance would have cost a practitioner more than random guessing. This is detailed in Section 3.5.
  • Universal failure on NG_009 (Ethiopian Airlines v Polaris Bank) – all six models missed the statute-bar point, demonstrating systematic procedural blindness. This is the most significant failure in the Nigerian dataset. Even with full citation context, only ChatGPT and Gemini corrected their predictions.
  • Universal failure on NG_020 (Atiba v Suberu) – all six models predicted the borrower's appeal would succeed (0/6), missing the parol evidence trap. This is equally significant as NG_009, revealing a systematic weakness in document hierarchy reasoning. With clues, five of six corrected (all except DeepSeek).
  • Translation Framework corrected three scores – without it, NG_013 DeepSeek, NG_019 DeepSeek, and NG_019 Perplexity would have been incorrectly scored, affecting final rankings.

Industrial Implication for Nigerian Practice:

ChatGPT with citations is procurement-ready for primary legal research (100% accuracy). DeepSeek must be placed on restricted vendor list pending independent verification protocol due to its 50% unaided accuracy and negative Expected Value pattern (EV approaching -1.0 on multiple cases).

3.5 DeepSeek Risk Signal – Nigeria

DeepSeek's performance in Nigeria exhibits the same negative Expected Value pattern observed in the US dataset. In several With-Clue instances, DeepSeek expressed high confidence on incorrect predictions, generating:

  • Expected Value penalties that would have cost a practitioner more than random guessing
  • Brier Scores approaching 1.0 on overconfident errors (e.g., NG_005, NG_006, NG_011)
  • Zero correction on NG_020 even with full citation context, unlike all other models

Implication for Nigerian practitioners: DeepSeek should not be used for Nigerian legal research without extensive independent verification. Its combination of low accuracy and high confidence on errors creates a "worst of both worlds" risk profile.

Comparative African Governance Evidence:

This risk pattern is not isolated. In South Africa, the Mavundla v MEC [2025] ZAKZPHC 2 case saw a legal team submit seven hallucinated citations – demonstrating that even Africa's most developed legal system is vulnerable. In Kenya, Republic v Public Procurement Administrative Review Board [2026] KEHC 1620 involved counsel relying on AI-generated legal provisions that did not exist. Rwanda offers a constructive counter-model through its National AI Policy 2023, which mandates transparency for AI in public sector decisions. Nigeria currently lacks any binding AI governance law, placing the entire burden of risk mitigation on individual firms.

SECTION 4: UNITED KINGDOM – COMPLETE FINDINGS

4.1 Without Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
UK_001Allowed111111
UK_002Allowed101111
UK_003Allowed110111
UK_004Allowed111111
UK_005Allowed011110
UK_006Dismissed110011
UK_007Allowed111110
UK_008Dismissed100111
UK_009Allowed111111
UK_010Dismissed111111
UK_011Allowed001111
UK_012Dismissed011110
UK_013Allowed011010
UK_014Allowed100110
UK_015Allowed111111
UK_016Dismissed111110
UK_017Allowed110111
UK_018Dismissed111111
UK_019Allowed101110
UK_020Dismissed111110
TOTAL12A/8D161515182012

4.2 With Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
UK_001Allowed111111
UK_002Allowed101111
UK_003Allowed111111
UK_004Allowed011001
UK_005Allowed111011
UK_006Dismissed101111
UK_007Allowed111111
UK_008Dismissed111111
UK_009Allowed111011
UK_010Dismissed111111
UK_011Allowed111111
UK_012Dismissed111111
UK_013Allowed111011
UK_014Allowed110101
UK_015Allowed011100
UK_016Dismissed111111
UK_017Allowed111111
UK_018Dismissed111111
UK_019Allowed111111
UK_020Dismissed111111
TOTAL12A/8D181819161719

4.3 UK: Summary Statistics

ModelNo Clue ScoreNo Clue AccuracyWith Clue ScoreWith Clue AccuracyImprovement
ChatGPT16/2080.0%18/2090.0%+2
Gemini15/2075.0%18/2090.0%+3
Claude15/2075.0%19/2095.0%+4
Grok18/2090.0%16/2080.0%-2
DeepSeek20/20100.0%17/2085.0%-3
Perplexity12/2060.0%19/2095.0%+7

4.4 UK: Key Findings

  • DeepSeek achieves perfect score without clues (20/20) – unique among all models across all jurisdictions. This demonstrates exceptional UK legal knowledge in its training data, including correct prediction of 2026 judgments.
  • Perplexity shows the largest improvement (+7 cases) – from last place (60%) to joint first (95%) with clues, confirming it is a retrieval-first model that performs best when citations are provided.
  • Claude achieves joint highest With-Clue score (19/20, 95%) with strong improvement (+4), indicating comprehensive UK case law coverage.
  • Grok declines with clues (-2) – the only model to perform worse with citation context, suggesting citation retrieval may introduce noise rather than signal. This is the second jurisdiction where Grok declines.
  • DeepSeek regresses with clues (-3) – from perfect score to 17/20, demonstrating that citation clues can override sound doctrinal reasoning with retrieved-but-inaccurate knowledge. DeepSeek expressed 100% confidence on two of its three With-Clue errors (UK_004 and UK_015).
  • Seven cases were universally correct across all models in both conditions (UK_001, UK_007, UK_010, UK_016, UK_017, UK_018, UK_020).

Industrial Implication for UK Practice:

DeepSeek's perfect unaided score (20/20) is deceptive – firms should require mandatory citation verification for any DeepSeek-assisted work due to its regression pattern with clues (-3) and overconfidence on errors (100% confidence on UK_004 and UK_015). For UK law with citations, Claude (95%) and Perplexity (95%) are safer choices.

SECTION 5: UNITED STATES – COMPLETE FINDINGS

5.1 Without Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
US_001Allowed011100
US_002Dismissed010101
US_003Dismissed010110
US_004Dismissed011110
US_005Allowed011110
US_006Dismissed111111
US_007Allowed111011
US_008Dismissed111101
US_009Allowed111111
US_010Dismissed111101
US_011Allowed011100
US_012Dismissed111101
US_013Dismissed111001
US_014Allowed111101
US_015Dismissed011100
US_016Allowed111111
US_017Dismissed111101
US_018Allowed111101
US_019Dismissed111110
US_020Allowed111111
TOTAL10A/10D13201818913

5.2 With Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
US_001Allowed010100
US_002Dismissed111101
US_003Dismissed110110
US_004Dismissed111110
US_005Allowed010111
US_006Dismissed111111
US_007Allowed111110
US_008Dismissed111111
US_009Allowed011111
US_010Dismissed111111
US_011Allowed011101
US_012Dismissed111101
US_013Dismissed010100
US_014Allowed010101
US_015Dismissed111111
US_016Allowed111111
US_017Dismissed111101
US_018Allowed111101
US_019Dismissed111110
US_020Allowed111111
TOTAL10A/10D142015201214

5.3 US: Summary Statistics

ModelNo Clue ScoreNo Clue AccuracyWith Clue ScoreWith Clue AccuracyImprovement
ChatGPT13/2065.0%14/2070.0%+1
Gemini20/20100.0%20/20100.0%0
Claude18/2090.0%15/2075.0%-3
Grok18/2090.0%20/20100.0%+2
DeepSeek9/2045.0%12/2060.0%+3
Perplexity13/2065.0%14/2070.0%+1

5.4 US: Key Findings

  • Gemini achieves perfect score in both conditions (20/20) – the only model to achieve perfect accuracy in both No-Clue and With-Clue in any jurisdiction. Its confidence calibration is exceptional (Avg Brier 0.0025 in No-Clue).
  • Grok achieves perfect score with clues (20/20), improving from 18/20 without clues – demonstrating strong US case law retrieval capability.
  • Claude regresses with clues (-3) – the only model to decline, suggesting citation-induced knowledge interference. This is a notable negative finding warranting further investigation.
  • DeepSeek performs near-random without clues (9/20, 45%) – below chance level on a balanced 10/10 split, confirming its training data is not US-focused.
  • US_013 (Medical Marijuana v Horn) was most difficult – only Gemini and Grok correct in With-Clue condition, demonstrating the frontier of recent (2025) jurisprudence. Even with citation, 4 of 6 models failed.
  • Cases where all models agreed – US_006 (Cuozzo) and US_009 (Becerra) were correctly predicted by all six models in both conditions.

Industrial Implication for US Practice:

Gemini is the only model that can be deployed without conditional prompting (100% accuracy both conditions). Claude requires citation-specific testing protocols due to its regression with clues (-3). For firms handling US-Nigeria cross-border work, ChatGPT with citations (100% on Nigeria, 70% on US) requires jurisdictional switching protocols.

SECTION 6: AUSTRALIA – COMPLETE FINDINGS

6.1 Without Clue – Complete Score Summary

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
AU_001Allowed111110
AU_002Dismissed000111
AU_003Allowed000101
AU_004Dismissed101101
AU_005Dismissed110101
AU_006Dismissed000101
AU_007Allowed011101
AU_008Dismissed111110
AU_009Allowed011101
AU_010Dismissed111101
AU_011Allowed111111
AU_012Dismissed010000
AU_013Allowed000000
AU_014Dismissed010000
AU_015Allowed111101
AU_016Dismissed111011
AU_017Allowed110111
AU_018Dismissed111111
AU_019Allowed111011
AU_020Dismissed111111
TOTAL—12151215915

6.2 With Clue – Complete Score Summary (CORRECTED)

Case IDGround TruthChatGPTGeminiClaudeGrokDeepSeekPerplexity
AU_001Allowed011100
AU_002Dismissed010111
AU_003Allowed111100
AU_004Dismissed111111
AU_005Dismissed101001
AU_006Dismissed011001
AU_007Allowed011011
AU_008Dismissed110111
AU_009Allowed111001
AU_010Dismissed111111
AU_011Allowed101011
AU_012Dismissed011111
AU_013Allowed110100
AU_014Dismissed011001
AU_015Allowed111111
AU_016Dismissed111111
AU_017Allowed111100
AU_018Dismissed111011
AU_019Allowed111000
AU_020Dismissed111111
TOTAL—141817121115

6.3 Australia: Summary Statistics (CORRECTED)

ModelNo Clue ScoreNo Clue AccuracyWith Clue ScoreWith Clue AccuracyImprovement
ChatGPT12/2060.0%14/2070.0%+2
Gemini15/2075.0%18/2090.0%+3
Claude12/2060.0%17/2085.0%+5
Grok15/2075.0%12/2060.0%-3
DeepSeek9/2045.0%11/2055.0%+2
Perplexity15/2075.0%15/2075.0%0

6.4 Australia: Key Findings (UPDATED)

  • Claude shows largest improvement (+5 cases) – from 12/20 to 17/20 with clues, demonstrating strong responsiveness to Australian citation context.
  • Gemini achieves highest With-Clue score (18/20, 90%) – most reliable for Australian law with citation context.
  • Grok declines with clues (-3) – now declines in three jurisdictions (UK -2, Australia -3, Nigeria 0), confirming a clear pattern of citation-induced performance degradation. This is the most significant model-specific negative finding.
  • AU_013 (Ecosse Property Holdings) produced universal failure (0/6) without clues – all models applied literal textual construction of "payable by the tenant" instead of purposive "de facto sale" interpretation. With clues, three models corrected (ChatGPT, Gemini, Grok).
  • AU_012 (Fischer v Nemeske) improved dramatically – from 1/6 (Gemini only) to 5/6 with clues, showing citation retrieval corrects trust law errors.
  • Three cases produced perfect 6/6 scores in No Clue condition: AU_011 (Simic), AU_018 (Westpac), and AU_020 (Productivity Partners). These were cases with well-established doctrinal principles well-represented in training data.
  • Partial verdict adjudications required – seven instances across both conditions, with Hard Boundary Rule applied consistently.
  • Correction Note: The AU_018 Grok (With Clue) score was corrected from 1 → 0, reducing Grok's With-Clue total from 13 to 12 and its accuracy from 65% to 60%. This correction has been applied consistently throughout the report.

Industrial Implication for Australian Practice:

Grok should be deployed with a mandatory "Partial Clue" protocol (year + court level only, no full citations) to optimise its 90% accuracy potential (discovered in Experiment 1, Section EXP-1). Without this protocol, Grok's accuracy falls to 65% with full citations. Gemini remains the most reliable for Australian law with citations (90%).

Sections 8–9 · Model specialisation by jurisdiction and Commonwealth lineageReport lines 530–577

SECTION 8: JURISDICTIONAL BIAS ANALYSIS

8.1 Model-Specialization Patterns (No Clue)

ModelBest JurisdictionWorst JurisdictionGapInterpretation
ChatGPTUK (80%)Australia (60%)20%Balanced
GeminiUS (100%)Nigeria (60%)40%Strong US bias
ClaudeUS (90%)Australia (60%)30%Moderate US bias
GrokUK/US (90%)Australia (75%)15%Balanced
DeepSeekUK (100%)US/Australia (45%)55%Extreme UK bias
PerplexityNigeria/Australia (75%)UK (60%)15%Slight Commonwealth preference

8.2 Key Findings on Jurisdictional Bias

  • Gemini shows strongest US bias – 40% gap between US (100%) and Nigeria (60%). This suggests US-centric training data with limited Nigerian legal content.
  • DeepSeek shows extreme UK bias – 55% gap between UK (100%) and US/Australia (45%). This is the most dramatic jurisdictional bias observed.
  • ChatGPT is most jurisdiction-agnostic – range of only 20% across all four jurisdictions, making it the most balanced model.
  • Perplexity performs best on Nigeria and Australia – the opposite of expected bias pattern, suggesting different training data composition.
  • Grok shows balanced performance – only 15% range, making it the second most jurisdiction-independent model after ChatGPT.
  • The "US/UK vs Nigeria/Australia" hypothesis is rejected – patterns are model-specific, not uniform. Some models (DeepSeek) show extreme UK bias, others (Perplexity) show reverse bias.

8.3 Bias Mitigation Procurement Protocol:

For IT managers selecting LLM vendors for Nigerian legal practice:

  • Vendor Disclosure Requirement – Require jurisdiction-specific accuracy claims to be verified against independent benchmarks (this study). Vendors must disclose training data composition by jurisdiction.
  • Bias Testing Mandate – Include jurisdiction sensitivity testing (Chi-square as per Section 2 Part 2) in vendor evaluation. A statistically significant gap (>0.05 Chi-square p-value) between US and Nigerian performance triggers enhanced verification requirements.
  • Contractual Clause – Material performance degradation (>15% gap between advertised and actual Nigerian accuracy) constitutes breach, allowing firm to terminate or demand remediation.
  • Ongoing Monitoring – Quarterly re-benchmarking on held-out Nigerian cases from the master dataset (20 cases available upon request). Maintain restricted vendor list for models exhibiting negative Expected Value.
  • Example Application: DeepSeek's 40-55% jurisdictional gaps and negative EV trigger automatic placement on restricted list, requiring case-by-case partner approval for any use.

SECTION 9: COMMONWEALTH LINEAGE ANALYSIS

9.1 UK → Australia → Nigeria Trend (No Clue)

ModelUKAustraliaNigeriaTrendSignificance
ChatGPT80%60%70%MixedNo clear trend
Gemini75%75%60%↓ Nigeria onlyUK→AU stable, NG drop
Claude75%60%70%MixedNo clear trend
Grok90%75%80%MixedNo clear trend
DeepSeek100%45%50%↓↓ Sharp dropStrong lineage effect
Perplexity60%75%75%↑ IncreasingReverse lineage

9.2 Key Findings on Commonwealth Lineage

  • Only DeepSeek shows the expected Commonwealth degradation pattern – perfect UK knowledge collapses on Australia and Nigeria, suggesting training data hierarchy (UK → Commonwealth → rest).
  • Perplexity shows reverse trend – performs best on Nigeria and Australia, worst on UK, indicating different training data composition focused on developing common law jurisdictions.
  • Gemini shows stable performance across UK and Australia (both 75%), with drop only on Nigeria, suggesting specific gap in Nigerian legal knowledge rather than general Commonwealth degradation.
  • The Commonwealth lineage hypothesis is model-dependent – not a universal phenomenon; some models (ChatGPT, Claude, Grok) show no clear lineage effect.
  • This finding has implications for legal practitioners – models cannot be assumed to transfer knowledge across Commonwealth jurisdictions uniformly.
Section 10 · Performance by legal domainReport lines 578–632

SECTION 10: ANALYSIS BY LEGAL DOMAIN

10.1 Nigeria – Performance by Legal Domain

DomainCasesAvg Accuracy (NC)Risk LevelVerification RequirementNotable Findings
Contract Law (Labour)NG_001, 005, 00777.8%MEDIUMVerify reasoning (Type II risk)NG_001 "correct guessing"
Contract Law (Land/Mortgage)NG_004, 012, 013, 02058.3%HIGHMandatory manual verification of document hierarchyNG_020 universal failure
Tort LawNG_002, 008, 014-015, 01772.0%MEDIUMStandard verification of case lawStrong overall performance
Commercial (IP/Banking/Admiralty)NG_003, 009-011, 018-01966.7%HIGHMandatory procedural check for limitation periodsNG_009 universal failure
Criminal LawNG_016100%LOWRoutine verification acceptableAll models correct

Domain Insight: Contract law involving land/mortgage proved most challenging (58.3%), while tort law showed stronger performance (72.0%). Procedural cases (NG_009) and document hierarchy cases (NG_020) exposed systematic weaknesses.

10.2 UK – Performance by Legal Domain

DomainCasesAvg Accuracy (NC)Risk LevelVerification RequirementNotable Findings
IP/PatentUK_018100%LOWRoutine verificationUniversal correctness
Civil ProcedureUK_009100%LOWRoutine verificationUniversal correctness
Contract LawUK_002, 017, 02088.9%LOWRoutine verificationStrong performance
Banking/CommercialUK_003, 012, 01688.9%LOWRoutine verificationStrong performance
EmploymentUK_01583.3%MEDIUMStandard verificationContested in With-Clue
Tort LawUK_001, 004-006, 008, 010-01176.2%MEDIUMVerify recent (2026) case logicUK_004 (2026) contested
Property/EquityUK_013, 01458.3%HIGHMandatory manual verificationUK_013 difficult

Domain Insight: Property/equity cases proved most challenging (58.3%), while contract, banking, and IP showed strong performance. Scottish private law (UK_006) exposed knowledge gaps.

10.3 US – Performance by Legal Domain

DomainCasesAvg Accuracy (NC)Risk LevelVerification RequirementNotable Findings
Administrative LawUS_009100%LOWRoutine verificationUniversal correctness
Employment/ArbitrationUS_003, 014, 018, 02087.5%LOWRoutine verificationStrong performance
Banking/PreemptionUS_00783.3%MEDIUMStandard verificationModerate difficulty
Intellectual PropertyUS_001, 005-006, 008, 015-01779.8%MEDIUMVerify against modern statutesWell-documented doctrine
BankruptcyUS_010, 01275.0%MEDIUMStandard verificationMixed performance
Civil ProcedureUS_002, 004, 01166.7%HIGHManual check of jurisdiction rulesUS_011 (Mallory) difficult
RICO Civil LawUS_01350.0%HIGHMandatory manual verificationMost difficult (2025)

Domain Insight: RICO (US_013) proved most challenging (50%), reflecting recency effect (2025). Intellectual Property cases showed strong performance (79.8%), suggesting well-documented doctrine in training data.

10.4 Australia – Performance by Legal Domain (UPDATED)

DomainCasesAvg Accuracy (NC)Risk LevelVerification RequirementNotable Findings
Commercial LawAU_018-02088.9%LOWRoutine verificationStrong performance
Public International LawAU_00883.3%MEDIUMStandard verificationStrong performance
IP LawAU_004, 01775.0%MEDIUMStandard verificationStrong performance
Tort LawAU_009, 01566.7%HIGHManual check of duty of careModerate performance
Contract LawAU_001, 005, 010-011, 013-01463.9%HIGHManual check of purposive intentAU_013 universal failure
EquityAU_002, 012, 01658.3%HIGHMandatory manual verificationAU_012 difficult
Civil ProcedureAU_006, 00758.3%HIGHManual check of court rulesAU_006 difficult
Property LawAU_00350.0%HIGHMandatory manual verificationMixed performance

Domain Insight: Equity cases proved most challenging (58.3%), while commercial law showed strongest performance (88.9%). Contract interpretation cases (AU_013) revealed systematic weaknesses in purposive construction.

Domain Risk Map Summary:

For Nigerian practitioners, Criminal Law (Green/Low Risk) can be safely delegated to AI with routine verification. Tort and non-land contract (Yellow/Medium Risk) require standard verification including reasoning checks. Land/mortgage contract, commercial banking, and IP (Orange/High Risk) require mandatory manual verification. Procedural/limitation and document hierarchy (Red/Critical Risk) should never be delegated to AI – manual verification is mandatory regardless of model confidence.

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 4.1 · source-reported

Verdict accuracy by model and jurisdiction under No Clue (correct out of 20 per jurisdiction)

   Model        Nigeria      UK        US    Australia       Total   Accuracy

   ChatGPT           14       16        13           12      55/80       68.8%
   Gemini            12       15        20           15      62/80       77.5%
   Claude            14       15        18           12      59/80       73.8%
   Grok              16       18        18           15      67/80       83.8%
   DeepSeek          10       20         9            9      48/80       60.0%
   Perplexity        15       12        13           15      55/80       68.8%

   Pooled       81/120    96/120   91/120      78/120     346/480       72.1%

Thesis Table 4.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.2 · source-reported

Verdict accuracy by model and jurisdiction under With Clue (cor- rect out of 20 per jurisdiction)

      Model                Nigeria             UK             US    Australia            Total       Accuracy

      ChatGPT                      20             18           14                 14     66/80          82.5%
      Gemini                       19             18           20                 18     75/80          93.8%
      Claude                       17             19           15                 17     68/80          85.0%
      Grok                         16             16           20                 12     64/80          80.0%
      DeepSeek                     13             17           12                 11     53/80          66.3%
      Perplexity                   16             19           14                 15     64/80          80.0%

      Pooled              101/120       107/120        95/120         87/120           390/480         81.3%
         The Nigeria columns are the original No Clue (81/120, 67.5%) and With Clue (101/120,
                              84.2%) baselines of the Nigerian extension.

Thesis Table 4.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.3 · source-reported

Commonwealth-lineage comparison under No Clue (Jonckheere– Terpstra test of the ordering UK, Australia, Nigeria)

      Model          UK     Australia   Nigeria   Observed pattern            p

      ChatGPT      80.0%       60.0%      70.0%   Mixed                   0.5874
      Gemini       75.0%       75.0%      60.0%   Lower Nigerian result   0.4157
      Claude       75.0%       60.0%      70.0%   Mixed                   0.7861
      Grok         90.0%       75.0%      80.0%   Mixed                   0.5874
      DeepSeek     100.0%      45.0%      50.0%   Marked decline          0.0067
      Perplexity   60.0%       75.0%      75.0%   Reverse pattern         0.4157

Thesis Table 4.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.