Key takeaways
- Phase 4 is in progress.
- Researcher observations are not lawyer-verified findings.
- No Type II Logic Failure rates are reported as established results.
What was tested & why
Preliminary reasoning analysis
The researcher is extracting the ratio decidendi and lawyers are to review model justifications. This section documents the boundary rather than inventing a reasoning score.
Preliminary researcher observation; awaiting lawyer verification.
Source: Thesis §§3.12, 10.11 · source-reported unless otherwise noted.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 12 · The “correct guessing” observationReport lines 655–676
SECTION 12: THE "CORRECT GUESSING" PHENOMENON
12.1 Preliminary Evidence of Stochastic Parroting
| Case | Jurisdiction | Domain | Models Correct | Preliminary Reasoning Observation | Status |
|---|---|---|---|---|---|
| NG_001 | Nigeria | Labour | 5/6 | Models used "negligence" (general principle) instead of "statutory flavour" (jurisdiction-specific principle) | Phase 2 verification pending |
| AU_013 | Australia | Contract | 0/6 | All models applied literal construction, missing purposive interpretation | Confirmed universal failure |
| UK_013 | UK | Property | 3/6 NC, 5/6 WC | Mixed performance | Phase 2 verification pending |
12.2 Reasoning Classification Framework
In Phase 2 of this research, all model reasoning will be classified according to the following framework:
| Category | Definition | Example |
|---|---|---|
| Aligned | LLM reasoning matches the jurisdiction-specific ratio decidendi that actually decided the case | Perplexity in NG_001 (preliminary) |
| General-but-Correct | LLM applies general common law principle that produces the same outcome as the actual judgment (Type II Logic Failure if jurisdiction-specific principle exists) | ChatGPT, Gemini, Claude, Grok, DeepSeek in NG_001 (preliminary) |
| General-but-Wrong | LLM applies general principle that would produce wrong outcome (but verdict correct by coincidence) | (to be determined in Phase 2) |
| Misaligned | LLM reasoning contradicts or ignores the actual principle | All models in AU_013 |
12.3 Type II Logic Failure Definition
A Type II Logic Failure occurs when a model predicts the correct verdict (Score = 1) but its reasoning falls into the General-but-Correct category – it uses general common law principles rather than the jurisdiction-specific principle that actually decided the case. This is the signature of "stochastic parroting" – correct guessing without genuine understanding.
Preliminary Evidence: Based on initial analysis of NG_001, Type II Logic Failure may have occurred in 5 of 6 models (83%) for that case alone. Models predicted correctly using general negligence principles rather than the jurisdiction-specific "statutory flavour" principle that actually decided the case.
Significance (if confirmed): This would provide the first empirical support for Dahl et al.'s "stochastic parrot" hypothesis with real legal data – a major theoretical contribution demonstrating that LLMs may achieve correct outcomes through probabilistic guessing rather than genuine jurisdictional understanding. Full confirmation awaits Phase 2 reasoning analysis.
12.4 Implications for Legal Practice
- Correct verdict ≠ correct reasoning – practitioners cannot assume that accurate predictions reflect sound legal analysis. A model may be "correct" but for the wrong reasons.
- Verification of reasoning is essential – the Duty of Inquiry Checklist must include verification of the legal principles applied, not just the outcome.
- Citation context helps but does not guarantee correct reasoning – even with clues, models may retrieve correct outcomes without understanding the underlying ratio decidendi.
Phase 2 · Reasoning analysis (researcher key-reason review, awaiting lawyer verification)Report lines 2344–2454
PART 3: Phase 2 — Reasoning Analysis
Key Reason Verification (Step 1) - Researcher Analysis
Verdict & Reasoning Performance
(1) Nigeria With Clue Experiments
The following table presents the ground truth ratio decidendi for each of the 20 Nigerian cases, alongside the grade (P / F / N/A) and a concise justification for each of the six evaluated LLMs (ChatGPT, Gemini, Claude, Grok, DeepSeek, Perplexity).
- P = Correct verdict and correct reasoning (aligned with the ratio).
- F = Correct verdict but incorrect reasoning (Type II logic failure).
- N/A = Incorrect verdict (reasoning evaluation is moot).
reasoning; F = Correct verdict but wrong reasoning (Type II failure); N/A = Incorrect verdict.
| Case ID | Key Ratio (Abbreviated) | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|
| NG_001 | Master servant; only damages, no reinstatement | P | P | P | P | P | P |
| NG_002 | Vicarious liability – own operator negligent | P | P | P | P | P | P |
| NG_003 | Stay of execution – no special circumstances | F | F | F | F | F | F |
| NG_004 | Originating process = writ; waiver of objection | F | F | F | F | F | F |
| NG_005 | Second dismissal of non existent employee is nullity | P | P | P | N/A | N/A | N/A |
| NG_006 | Suspicion ≠ proof; concrete evidence required | F | N/A | N/A | N/A | N/A | F |
| NG_007 | Equity prevents benefiting from own wrong; 2 years’ salary | F | F | F | F | F | F |
| NG_008 | Conversion liability for paying wrong person | P | P | P | P | P | P |
| NG_009 | Limitation bar – suit filed after 6 years | F | F | N/A | N/A | N/A | N/A |
| NG_010 | Pre existing indebtedness survives breach | P | P | P | P | N/A | N/A |
| NG_011 | Admiralty ends at discharge; State High Court jurisdiction | P | P | P | N/A | N/A | P |
| NG_012 | 16 month delay defeats specific performance | P | P | P | P | P | P |
| NG_013 | Estoppel by conduct (s.169 Evidence Act) | P | P | P | P | P | P |
| NG_014 | Independent contractor – no vicarious liability | P | P | P | P | P | P |
| NG_015 | Presumption of negligence from failure to return goods | F | F | F | F | F | F |
| NG_016 | Prosecution need prove only ONE of four ingredients | F | F | F | F | F | F |
| NG_017 | Lack of publication defeats defamation | F | F | N/A | F | F | F |
| NG_018 | Court integrity prioritised over jurisdiction | F | F | F | F | N/A | F |
| NG_019 | Interest as of right for fiduciary breach | F | F | F | F | F | N/A |
| NG_020 | Parol evidence rule – deed supersedes letter | P | P | P | P | N/A | P |
Summary Table: Nigeria With Clue Experiments – Verdict & Reasoning Performance
| LLM | Correct Verdict (P+F) | Correct Reasoning (P) | Type II Failures (F) | Incorrect Verdict (N/A) | Reasoning Accuracy (of correct verdicts) | Overall Reasoning Accuracy (of all cases) |
|---|---|---|---|---|---|---|
| ChatGPT | 20 | 10 | 10 | 0 | 50.0% | 50.0% |
| Gemini | 19 | 10 | 9 | 1 | 52.6% | 50.0% |
| Claude | 17 | 10 | 7 | 3 | 58.8% | 50.0% |
| Grok | 16 | 8 | 8 | 4 | 50.0% | 40.0% |
| DeepSeek | 13 | 6 | 7 | 7 | 46.2% | 30.0% |
| Perplexity | 16 | 8 | 8 | 4 | 50.0% | 40.0% |
Notes:
- Correct Verdict = number of cases where the model’s predicted outcome matched the ground truth (P + F).
- Correct Reasoning = number of cases where the model’s legal reasoning aligned with the ratio decidendi (Grade P only).
- Type II Failure = correct verdict but wrong reasoning (Grade F).
- Incorrect Verdict = verdict did not match ground truth (Grade N/A).
- Reasoning Accuracy (of correct verdicts) = Correct Reasoning / Correct Verdict × 100.
- Overall Reasoning Accuracy = Correct Reasoning / 20 × 100.
Key Findings
- High Verdict Accuracy but Low Reasoning Fidelity
All models achieved high verdict correctness (ChatGPT 100%, Gemini 95%, Claude 85%, Grok 80%, Perplexity 80%, DeepSeek 65%). However, reasoning alignment with the true ratio decidendi was low: ChatGPT, Gemini, and Claude each had only 10 out of 20 cases with correct reasoning (50%), while Grok and Perplexity scored 40%, and DeepSeek 30%.
- Type II Logic Failures are Pervasive
Type II failures (correct verdict but wrong legal reasoning) occurred in 7–10 cases per model. This is the most dangerous error for legal AI because a correct outcome masks flawed logic, undermining trust and explainability.
- Strict Liability and Procedural Cases Caused Most Failures
Cases involving stay of execution (NG_003), procedural waiver (NG_004), equitable damages quantum (NG_007), presumption of negligence in bailment (NG_015), criminal burden “one of four ingredients” (NG_016), publication in defamation (NG_017), court integrity over jurisdiction (NG_018), and interest as of right for fiduciary breach (NG_019) were consistently mis reasoned by almost all models.
- DeepSeek’s Lower Verdict Accuracy Reflects Conservative Reasoning
DeepSeek had the fewest correct verdicts (13/20) but also exhibited different failure patterns. It correctly applied estoppel (NG_013) and vicarious liability (NG_002, NG_014) but struggled with limitation bars (NG_009) and admiralty jurisdiction (NG_011).
- No Model Achieved “Human Level” Reasoning
Even the best model (ChatGPT or Claude) only had correct reasoning in half of the cases, highlighting a substantial gap for research in legal reasoning evaluation.
Visualisation Image:
Figure: Reasoning & Verdict Performance - (1.) Nigeria With Clue
Bar chart description:
- Green = Correct Reasoning (P)
- Orange = Type II Failure (correct verdict, wrong reasoning)
- Red = Incorrect Verdict (N/A)
- Number on top of each bar = total correct verdicts (P+F).
Conclusion
The Nigeria With Clue experiments demonstrate that while large language models can often predict the correct legal outcome (verdict accuracy up to 100%), their legal reasoning – as measured against the true ratio decidendi – is alarmingly deficient. Half of all correct verdicts were supported by incorrect or irrelevant legal principles (Type II failures). This finding has critical implications for deploying LLMs in legal practice:
- Verdict alone is not a reliable metric for legal AI capability.
- Explainability and reasoning verification must be core components of any legal LLM system.
- Specific legal domains (procedural law, equitable remedies, strict liability nuances, defamation elements) are particularly challenging and require targeted fine tuning or retrieval augmented grounding.
Future work should focus on reasoning aware evaluation frameworks, adversarial testing of legal logic, and hybrid systems that combine LLM output with rule based verification against the ratio decidendi. The current results underscore that even the best models are not yet ready for unsupervised legal decision support.
APPENDICES:
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.