CrossLaw/Evidence library/Preliminary reasoning analysis
21 / 26 · Evidence libraryThesis §§3.12, 10.11

Reasoning-quality conclusions await lawyer verification.

Preliminary reasoning analysis

Key takeaways

  • Phase 4 is in progress.
  • Researcher observations are not lawyer-verified findings.
  • No Type II Logic Failure rates are reported as established results.

What was tested & why

Preliminary reasoning analysis

The researcher is extracting the ratio decidendi and lawyers are to review model justifications. This section documents the boundary rather than inventing a reasoning score.

Evidence boundary

Preliminary researcher observation; awaiting lawyer verification.

Source: Thesis §§3.12, 10.11 · source-reported unless otherwise noted.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Section 12 · The “correct guessing” observationReport lines 655–676

SECTION 12: THE "CORRECT GUESSING" PHENOMENON

12.1 Preliminary Evidence of Stochastic Parroting

CaseJurisdictionDomainModels CorrectPreliminary Reasoning ObservationStatus
NG_001NigeriaLabour5/6Models used "negligence" (general principle) instead of "statutory flavour" (jurisdiction-specific principle)Phase 2 verification pending
AU_013AustraliaContract0/6All models applied literal construction, missing purposive interpretationConfirmed universal failure
UK_013UKProperty3/6 NC, 5/6 WCMixed performancePhase 2 verification pending

12.2 Reasoning Classification Framework

In Phase 2 of this research, all model reasoning will be classified according to the following framework:

CategoryDefinitionExample
AlignedLLM reasoning matches the jurisdiction-specific ratio decidendi that actually decided the casePerplexity in NG_001 (preliminary)
General-but-CorrectLLM applies general common law principle that produces the same outcome as the actual judgment (Type II Logic Failure if jurisdiction-specific principle exists)ChatGPT, Gemini, Claude, Grok, DeepSeek in NG_001 (preliminary)
General-but-WrongLLM applies general principle that would produce wrong outcome (but verdict correct by coincidence)(to be determined in Phase 2)
MisalignedLLM reasoning contradicts or ignores the actual principleAll models in AU_013

12.3 Type II Logic Failure Definition

A Type II Logic Failure occurs when a model predicts the correct verdict (Score = 1) but its reasoning falls into the General-but-Correct category – it uses general common law principles rather than the jurisdiction-specific principle that actually decided the case. This is the signature of "stochastic parroting" – correct guessing without genuine understanding.

Preliminary Evidence: Based on initial analysis of NG_001, Type II Logic Failure may have occurred in 5 of 6 models (83%) for that case alone. Models predicted correctly using general negligence principles rather than the jurisdiction-specific "statutory flavour" principle that actually decided the case.

Significance (if confirmed): This would provide the first empirical support for Dahl et al.'s "stochastic parrot" hypothesis with real legal data – a major theoretical contribution demonstrating that LLMs may achieve correct outcomes through probabilistic guessing rather than genuine jurisdictional understanding. Full confirmation awaits Phase 2 reasoning analysis.

12.4 Implications for Legal Practice

  • Correct verdict ≠ correct reasoning – practitioners cannot assume that accurate predictions reflect sound legal analysis. A model may be "correct" but for the wrong reasons.
  • Verification of reasoning is essential – the Duty of Inquiry Checklist must include verification of the legal principles applied, not just the outcome.
  • Citation context helps but does not guarantee correct reasoning – even with clues, models may retrieve correct outcomes without understanding the underlying ratio decidendi.
Phase 2 · Reasoning analysis (researcher key-reason review, awaiting lawyer verification)Report lines 2344–2454

PART 3: Phase 2 — Reasoning Analysis

Key Reason Verification (Step 1) - Researcher Analysis

Verdict & Reasoning Performance

(1) Nigeria With Clue Experiments

The following table presents the ground truth ratio decidendi for each of the 20 Nigerian cases, alongside the grade (P / F / N/A) and a concise justification for each of the six evaluated LLMs (ChatGPT, Gemini, Claude, Grok, DeepSeek, Perplexity).

  • P = Correct verdict and correct reasoning (aligned with the ratio).
  • F = Correct verdict but incorrect reasoning (Type II logic failure).
  • N/A = Incorrect verdict (reasoning evaluation is moot).

reasoning; F = Correct verdict but wrong reasoning (Type II failure); N/A = Incorrect verdict.

Case IDKey Ratio (Abbreviated)ChatGPTGeminiClaudeGrokDeepSeekPerplexity
NG_001Master servant; only damages, no reinstatementPPPPPP
NG_002Vicarious liability – own operator negligentPPPPPP
NG_003Stay of execution – no special circumstancesFFFFFF
NG_004Originating process = writ; waiver of objectionFFFFFF
NG_005Second dismissal of non existent employee is nullityPPPN/AN/AN/A
NG_006Suspicion ≠ proof; concrete evidence requiredFN/AN/AN/AN/AF
NG_007Equity prevents benefiting from own wrong; 2 years’ salaryFFFFFF
NG_008Conversion liability for paying wrong personPPPPPP
NG_009Limitation bar – suit filed after 6 yearsFFN/AN/AN/AN/A
NG_010Pre existing indebtedness survives breachPPPPN/AN/A
NG_011Admiralty ends at discharge; State High Court jurisdictionPPPN/AN/AP
NG_01216 month delay defeats specific performancePPPPPP
NG_013Estoppel by conduct (s.169 Evidence Act)PPPPPP
NG_014Independent contractor – no vicarious liabilityPPPPPP
NG_015Presumption of negligence from failure to return goodsFFFFFF
NG_016Prosecution need prove only ONE of four ingredientsFFFFFF
NG_017Lack of publication defeats defamationFFN/AFFF
NG_018Court integrity prioritised over jurisdictionFFFFN/AF
NG_019Interest as of right for fiduciary breachFFFFFN/A
NG_020Parol evidence rule – deed supersedes letterPPPPN/AP

Summary Table: Nigeria With Clue Experiments – Verdict & Reasoning Performance

LLMCorrect Verdict (P+F)Correct Reasoning (P)Type II Failures (F)Incorrect Verdict (N/A)Reasoning Accuracy (of correct verdicts)Overall Reasoning Accuracy (of all cases)
ChatGPT201010050.0%50.0%
Gemini19109152.6%50.0%
Claude17107358.8%50.0%
Grok1688450.0%40.0%
DeepSeek1367746.2%30.0%
Perplexity1688450.0%40.0%

Notes:

  • Correct Verdict = number of cases where the model’s predicted outcome matched the ground truth (P + F).
  • Correct Reasoning = number of cases where the model’s legal reasoning aligned with the ratio decidendi (Grade P only).
  • Type II Failure = correct verdict but wrong reasoning (Grade F).
  • Incorrect Verdict = verdict did not match ground truth (Grade N/A).
  • Reasoning Accuracy (of correct verdicts) = Correct Reasoning / Correct Verdict × 100.
  • Overall Reasoning Accuracy = Correct Reasoning / 20 × 100.

Key Findings

  • High Verdict Accuracy but Low Reasoning Fidelity

All models achieved high verdict correctness (ChatGPT 100%, Gemini 95%, Claude 85%, Grok 80%, Perplexity 80%, DeepSeek 65%). However, reasoning alignment with the true ratio decidendi was low: ChatGPT, Gemini, and Claude each had only 10 out of 20 cases with correct reasoning (50%), while Grok and Perplexity scored 40%, and DeepSeek 30%.

  • Type II Logic Failures are Pervasive

Type II failures (correct verdict but wrong legal reasoning) occurred in 7–10 cases per model. This is the most dangerous error for legal AI because a correct outcome masks flawed logic, undermining trust and explainability.

  • Strict Liability and Procedural Cases Caused Most Failures

Cases involving stay of execution (NG_003), procedural waiver (NG_004), equitable damages quantum (NG_007), presumption of negligence in bailment (NG_015), criminal burden “one of four ingredients” (NG_016), publication in defamation (NG_017), court integrity over jurisdiction (NG_018), and interest as of right for fiduciary breach (NG_019) were consistently mis reasoned by almost all models.

  • DeepSeek’s Lower Verdict Accuracy Reflects Conservative Reasoning

DeepSeek had the fewest correct verdicts (13/20) but also exhibited different failure patterns. It correctly applied estoppel (NG_013) and vicarious liability (NG_002, NG_014) but struggled with limitation bars (NG_009) and admiralty jurisdiction (NG_011).

  • No Model Achieved “Human Level” Reasoning

Even the best model (ChatGPT or Claude) only had correct reasoning in half of the cases, highlighting a substantial gap for research in legal reasoning evaluation.

Visualisation Image:

Figure: Reasoning & Verdict Performance - (1.) Nigeria With Clue

Bar chart description:

  • Green = Correct Reasoning (P)
  • Orange = Type II Failure (correct verdict, wrong reasoning)
  • Red = Incorrect Verdict (N/A)
  • Number on top of each bar = total correct verdicts (P+F).

Conclusion

The Nigeria With Clue experiments demonstrate that while large language models can often predict the correct legal outcome (verdict accuracy up to 100%), their legal reasoning – as measured against the true ratio decidendi – is alarmingly deficient. Half of all correct verdicts were supported by incorrect or irrelevant legal principles (Type II failures). This finding has critical implications for deploying LLMs in legal practice:

  • Verdict alone is not a reliable metric for legal AI capability.
  • Explainability and reasoning verification must be core components of any legal LLM system.
  • Specific legal domains (procedural law, equitable remedies, strict liability nuances, defamation elements) are particularly challenging and require targeted fine tuning or retrieval augmented grounding.

Future work should focus on reasoning aware evaluation frameworks, adversarial testing of legal logic, and hybrid systems that combine LLM output with rule based verification against the ratio decidendi. The current results underscore that even the best models are not yet ready for unsupervised legal decision support.

APPENDICES:

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.