CrossLaw/Evidence library/Appendices
22 / 26 · Evidence libraryThesis Chapter 12; historical benchmarking reports

Case records and report-only research questions behind the narrative.

Appendices

Key takeaways

  • Appendix tables retain their thesis numbering.
  • Table 12.6 lists benchmark cases and ground truth.
  • The earlier reports’ 35 open questions are reproduced separately from the thesis’s 11 RQs.

What was tested & why

Appendices

Thesis tables appear below. The newly supplied historical reports also contain a 35-question follow-up catalogue, reproduced separately here as report-only material. Their per-case Part I score sheets, preliminary P/F/N/A reasoning grades and supplementary outputs remain outside the verified thesis evidence set.

Evidence boundary

Report-only questions are not answers or established findings; reasoning grades await lawyer verification.

Source: Thesis Chapter 12; historical benchmarking reports · source-reported unless otherwise noted.

Historical benchmarking and weekly reports · report-only

35 open questions, not 35 thesis RQs

The earlier reports include this catalogue of possible follow-up questions. Some prompts presuppose explanations not established by the final thesis, and items 5/27 and 14/28 overlap. They are reproduced as questions, not answered findings, endorsed hypotheses, or the thesis’s 11 research questions.

01

Why does Grok consistently decline with citation clues? Is its training data structured differently such that citation retrieval introduces noise rather than signal?

02

What explains DeepSeek’s extreme UK specialization? Would it perform well on other Commonwealth jurisdictions (e.g., Canada, New Zealand, India)?

03

Why does Claude regress with clues in the US but improve elsewhere?

04

What drives Perplexity’s dramatic clue effect? Does it rely more heavily on retrieval than reasoning?

05

Are there ‘families’ of models with similar behavior patterns?

06

Why do procedural/limitation cases cause universal failure?

07

Why does document hierarchy (parol evidence) cause universal failure?

08

Why does purposive construction cause universal failure?

09

Which legal domains are most and least reliably handled?

10

How do models perform on constitutional vs. private law questions?

11

Do models perform better on landmark vs. obscure cases?

12

What is the precise training data cutoff for each model?

13

How does case age interact with jurisdiction?

14

How will these findings change as models update?

15

Why is DeepSeek’s calibration perfect in UK but poor elsewhere?

16

What causes the ‘overconfidence trap’ in some models?

17

Do models know when they don’t know? Which models have the best ‘metacognitive awareness’?

18

How much did the Translation Framework affect rankings?

19

How often does the Hard Boundary Rule apply?

20

What is the optimal fact pattern length?

21

Can we create a jurisdiction-specific reliability index?

22

What is the cost-benefit of using multiple models?

23

Can we predict case difficulty for LLMs?

24

How should the Duty of Inquiry Checklist be customized by jurisdiction?

25

How do these models compare to human lawyers?

26

How do these models compare to each other statistically?

27

Are there ‘families’ of models with similar behavior?

28

How will these findings change as models update?

29

Can fine-tuning on Nigerian law improve performance?

30

What does ‘understanding’ mean in the context of LLM legal reasoning?

31

Is jurisdictional bias a form of epistemic injustice?

32

What does the clue effect reveal about model architecture?

33

Can we automate the Duty of Inquiry Checklist?

34

How should law firms integrate these findings into practice?

35

What is the economic impact of using optimal vs. suboptimal models?

Weekly benchmarking report, “New Research Questions Emerging from the Findings,” pp. 40–41. Report-only; the thesis §1.3 and §10.9 supply the authoritative 11 RQs and answers.

Supplementary records held apart

The reports also contain per-case Part I score matrices, Nigeria With Clue researcher P/F/N/A reasoning grades, and supplementary statistical outputs. These have not been independently reconciled cell by cell against the final thesis. The P/F/N/A entries are not lawyer-verified reasoning-quality findings (thesis §3.12); no Type II Logic Failure rate is inferred from them.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Executive summary, headline findings and timelineReport lines 15–48

EXECUTIVE SUMMARY

This report presents the complete, verified findings from a comprehensive benchmarking study evaluating six leading large language models (ChatGPT, Gemini, Claude, Grok, DeepSeek, Perplexity) on their ability to predict appellate court outcomes across four common law jurisdictions: Nigeria, United Kingdom, United States, and Australia.

1. The Master Research Timeline (For your Presentation Slide)

This is the high-level overview perfect for Slide 10 of your presentation.

YEAR 1: FOUNDATION & EXPERIMENTATION | Phase | Activity | Status | | :--- | :--- | :--- | | Literature & design | Spring 2025 – Ongoing | ✅ Continuous (50+ papers reviewed & 20+ in-depth study) | | Case selection & extraction | Summer 2025/2026 – Autumn 2026 | ✅ Complete | | LLM testing (960 experiments) | Autumn 2026 | ✅ Complete | | Scoring & descriptive analysis | Autumn 2026 | ✅ Complete | | Statistical analysis (Friedman, McNemar, etc.) | Autumn 2026 | ✅ Complete | | Adaptive follow up experiments (addressing emerged patterns) | Autumn 2026 | Ongoing |

YEAR 2: VERIFICATION & TRANSLATION | Phase | Activity | Status | | :--- | :--- | :--- | | Reasoning verification (Phase 2 – lawyer HITL) | Autumn 2026 / Spring 2026 | Planned | | All Industrial toolkit Components Development | Spring 2026 | Planned | | Website/App Development | Summer 2026/2027 – Autumn 2027 | Planned | | Research thesis | Summer 2026/2027 – 2027 | Planned |

2. The Detailed Task Breakdown (For Project Planning & Speaker Notes)

This breakdown details the realistic allocation of the remaining phases mapped across the WSU academic calendar.

AUTUMN 2026 (Finalizing Experimental Work & Analysis) Current Semester – Finalizing the data | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phases 1–3 | Finalizing Adaptive Follow-up Experiments (Grok, DeepSeek, ChatGPT) | 3–4 weeks (Concurrent) | | Phase 4 | Researcher analysis of KeyReason (Ratio decidendi) & Further Literature Study | 4 weeks |

MID-YEAR BREAK / EARLY SPRING 2026 (Reasoning Verification) June/July/August – Allowing maximum flexibility for external partners | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 5 | Lawyer verification of KeyReason (Phase 2 – Lawyer HITL) | 4–6 weeks (Extensive buffer for practitioner schedules) | | Phase 6 | Type II Logic Failure quantification | 1 week |

SPRING 2026 (Industrial Translation) August to November – Core development phase | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 7 | All Industrial Toolkit Components Development (8 outputs) | 2+ weeks | | Phase 8 | Practitioner Manual drafting | 2+ weeks |

SUMMER 2026/2027 ONWARDS (Thesis Finalisation) December to mid-2027 – Final synthesis and submission | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 9 | Core thesis writing, comprehensive results synthesis, and industrial toolkit integration | Ongoing (3+ months) | | Phase 10 | Final thesis formatting, defense preparation, review, and official submission | 3–4 weeks |

Total Observations: 960 (fully verified and audited)

Important Note on Statistical Significance

The findings presented in this Executive Summary are descriptive — they reflect observed patterns in the data. While the dataset is comprehensive (960 observations), statistical testing (planned for Phase 1 completion) will determine which observed differences are statistically significant given the sample size of 20 cases per jurisdiction. The descriptive findings provide strong directional evidence and actionable insights, but final conclusions about the robustness of specific comparisons await statistical confirmation.

Headline Findings (Descriptive)

  • No single model dominates all jurisdictions – Performance is highly model-dependent and jurisdiction-dependent.
  • Jurisdictional bias is real but model-specific – Gemini shows strong US bias; DeepSeek shows extreme UK specialization; ChatGPT and Perplexity are more balanced.
  • Clue effect is model-dependent – Most models improve with citation context (+9.1% average), but Grok declines (-3), revealing critical differences in how models use retrieval vs. reasoning.
  • "Correct guessing" phenomenon observed – Preliminary analysis suggests models often predict correctly using general principles rather than jurisdiction-specific ratio decidendi (e.g., NG_001: 5/6 correct but used "negligence" instead of "statutory flavour"). This will be systematically verified in Phase 2 reasoning analysis.
  • Universal failures reveal systematic gaps – Cases involving procedural law, limitation periods, statutory interpretation, and document hierarchy consistently challenge all models (NG_009: 0/6; NG_020: 0/6; AU_013: 0/6).
  • DeepSeek is uniquely UK-specialized – Perfect 20/20 on UK without clues, but near-random on US (9/20) and Australia (9/20).
  • Gemini is the most reliable with citation context – 93.8% average accuracy across all four jurisdictions with clues.
  • Grok is the most consistent without clues – 83.8% average accuracy across all four jurisdictions without any case identification.

Industrial Translation:

This research further contributes eight industrial governance outputs including a Traceable Accountability three pillar framework, risk classified model selection matrix, domain specific risk map, five phase implementation roadmap, vendor selection scorecard, expanded Duty of Inquiry Checklist, AI Governance Committee charter template, and incident tracking protocol – translating 960 empirical observations into deployable IT management tools for the Nigerian legal sector.

Sections 15–19 · Theoretical contributions, statistical plan, Phase 2, emerging questions and conclusionReport lines 858–1108

SECTION 15: THEORETICAL CONTRIBUTIONS

15.1 Empirical Evidence for "Stochastic Parrot" Hypothesis (Preliminary)

This research offers preliminary empirical support for Dahl et al.'s "stochastic parrot" hypothesis using real legal data. The NG_001 finding – 5/6 models correct but apparently using the wrong reasoning principle – suggests that LLMs may achieve correct outcomes through probabilistic guessing rather than genuine understanding. Full confirmation awaits Phase 2 reasoning analysis.

15.2 Type II Logic Failure as a Metric

You have operationalized a new metric: Type II Logic Failure (correct verdict, wrong reasoning). This distinguishes between models that understand law (Aligned) and models that successfully guess (General-but-Correct) – a crucial distinction for legal practitioners.

15.3 Model-Specific Jurisdictional Bias Taxonomy

You have created the first taxonomy of LLM jurisdictional bias, demonstrating that:

  • Gemini = US-specialized
  • DeepSeek = UK-specialized
  • ChatGPT = Jurisdiction-agnostic
  • Perplexity = Commonwealth-preferring (Nigeria/Australia)
  • Grok = Most balanced across jurisdictions
  • Claude = Moderate US bias with strong UK retrieval

15.4 Translation Framework Methodology

Your Translation Framework for converting substantive liability predictions to procedural appellate outcomes is a methodological innovation that can be adopted by future legal AI benchmarking studies. It ensures fair scoring regardless of output vocabulary.

15.5 Hard Boundary Rule

The Hard Boundary Rule provides a principled method for handling internally inconsistent responses, ensuring scoring consistency and preventing models from benefiting from contradictions.

15.6 Recency Effect Quantification

You have quantified the recency effect in legal LLM performance – a 16.8% accuracy drop for 2025 cases compared to 2015 cases, providing empirical evidence of training data limitations.

15.7 Domain-Specific Performance Mapping

You have created the first detailed mapping of LLM performance across legal domains, identifying systematic weaknesses in procedural law, purposive construction, document hierarchy, and equity.

15.8 Citation Interference Pattern

You have identified a novel phenomenon: citation interference – where providing citation context actively degrades model performance. Grok's consistent decline with clues (-3 overall) is the first documented case of this pattern in legal LLM research.

15.9 Industrial Translation Framework

This research contributes a methodology for translating academic benchmark findings into industrial governance tools – a meta-contribution that enables future researchers to replicate the translation process.

Academic FindingIndustrial OutputTranslation MethodEvidence Section
Model accuracy tables (Sections 3-7)Risk-classified selection matrix (14.1)Weighted scoring: Nigerian accuracy 30%, cross-jurisdictional consistency 20%, high-confidence error rate 15%7.1-7.2
Type II Logic Failure (12.3)Duty of Inquiry Checklist (14.4)Extracted verification items from failure case analysis (NG_001, AU_013)12.3
Citation interference pattern (Grok)Conditional prompting protocol (14.2.5)Tested across 5 conditions to identify optimal input formatExperiment 1
Recency effect (16.8% drop, 11.2)"Post-2023" verification rule (14.3)Operational guardrail with clear implementation threshold11.2
Procedural blindness (NG_009)Mandatory manual verification for limitation questionsDomain-specific risk classification (Critical)13.1
Negative Expected Value (DeepSeek)"NOT RECOMMENDED" procurement designationEV threshold: < -0.5 triggers automatic restricted status3.5, 5.3
RLHF-induced agreeableness (Exp 3)Adversarial pushback protocol (14.2.5)Two-step verification: fresh rerun + challengeExperiment 3
Shadow IT risk (31%)Phase 1 implementation step (14.6)Mandatory audit before policy adoption14.6

This framework ensures that every empirical finding generates an actionable output for legal practitioners and IT managers.

SECTION 16: NEXT STEPS – STATISTICAL TESTING

16.1 Research Questions and Statistical Confirmation (UPDATED)

Research QuestionInitial Finding from Descriptive DataPrimary Statistical Test to Confirm
RQ1: Which model is most jurisdiction-independent?Grok leads without clues (83.8%); Gemini leads with clues (93.8%)Friedman Test + Kendall's W – to determine if differences across jurisdictions are statistically significant
RQ1a: Do models show jurisdiction sensitivity?Yes – DeepSeek: UK 100% vs US 45%; Gemini: US 100% vs Nigeria 60%Friedman Test by Model – to test if within-model variation across jurisdictions is significant
RQ1b: Which model has most consistent accuracy?ChatGPT: 60-80% range; Grok: 75-90% rangeCoefficient of Variation – to quantify consistency relative to mean accuracy
RQ2: Does adding clues improve accuracy?Average improvement +9.1%; Grok declines (-3); Perplexity improves most (+7 in UK)McNemar's Test – to determine if improvement for each model is statistically significant
Narrative A: US vs Nigeria biasGemini gap: 40% (US 100% vs Nigeria 60%)Mann-Whitney U + Cliff's Delta – to test if US-Nigeria gap is significant and measure effect size
Narrative B: Commonwealth lineageDeepSeek shows sharp drop (100% UK → 45% AU); others show mixed trendsJonckheere-Terpstra Test – to test for ordered trend across UK→AU→NG
Calibration: Are models appropriately confident?DeepSeek: perfect NC calibration but degraded WC; Gemini: excellent both conditionsReliability Diagrams + ECE – to visualize and quantify calibration quality

16.2 Why Statistical Testing is Necessary

Descriptive FindingWhat Statistics Will Confirm
"Grok scores 18/20 in UK without clues"Is this significantly better than chance given the 60/40 case distribution?
"Gemini shows 40% US-Nigeria gap"Is this gap statistically significant with n=20 per jurisdiction?
"Perplexity improves by +7 in UK"Is this improvement significant given the sample size?
"DeepSeek's UK specialization"Is the difference between UK and other jurisdictions significant?
"Claude regresses with clues in US"Is this regression statistically meaningful or random variation?

16.3 Complete Statistical Analysis Workflow

StepAnalysisTestOutputWhat It Confirms
1Descriptive statisticsMean, Std, CVBaseline metricsInitial performance patterns
2Overall model comparisonFriedmanp-value for overall differenceWhether models differ significantly overall
3If significantWilcoxon + BonferroniWhich models differSpecific pairwise differences
4Clue vs. No ClueMcNemarp-value per modelWhether clue effect is significant
5US vs. Nigeria biasMann-Whitney + Cliff's Deltap-value + effect sizeWhether bias gap is significant and its magnitude
6Commonwealth trendJonckheere-Terpstrap-value for ordered trendWhether lineage effect exists
7Confidence calibrationReliability diagrams + ECEVisual + numericCalibration quality
8Effect sizesCliff's Delta, r, Kendall's WMagnitude of differencesPractical significance
9Correct Guessing AnalysisResearcher + LawyerReasoning ClassificationType II Logic Failure rates

16.4 Timeline for Statistical Testing

PhaseActivityEstimated Duration
1Data preparation for statistical analysis1 day
2Descriptive statistics and visualization1 day
3Friedman tests (overall and by model)1 day
4Pairwise Wilcoxon with Bonferroni1 day
5McNemar tests for clue effect1 day
6Mann-Whitney U for US/Nigeria bias1 day
7Jonckheere-Terpstra for Commonwealth trend1 day
8Reliability diagrams and ECE calculation1 day
9Effect size calculations and synthesis1 day
10Results interpretation and reporting2 days

Total Estimated Time: 10-12 days

SECTION 17: PHASE 2 – REASONING ANALYSIS

17.1 Two-Step KeyReason Verification

Step 1: Researcher Analysis

  • Study each case file to identify ratio decidendi
  • Document with page/paragraph references
  • Preliminary classification for each LLM using the four-category framework
  • Flag potential Type II Logic Failures (General-but-Correct)

Step 2: Lawyer Verification

  • Compile complete package for each case:
  • Original judgment PDF
  • Researcher's extracted ratio decidendi
  • Each LLM's stated reasoning
  • Preliminary classification
  • Legal practitioner provides final verification
  • Update master sheet with verified Reasoning Classification

17.2 Reasoning Classification Framework

CategoryDefinitionScoring Impact
AlignedLLM reasoning matches jurisdiction-specific principleGenuine understanding demonstrated
General-but-CorrectLLM applies general common law principle that produces same outcome (Type II Logic Failure if jurisdiction-specific principle exists)Correct outcome but wrong reasoning – signals stochastic parroting
General-but-WrongLLM applies general principle that would produce wrong outcome (but verdict correct by coincidence)Correct by luck – unreliable for future cases
MisalignedLLM reasoning contradicts or ignores the actual principleFundamental error

17.3 Cases Selected for Initial Reasoning Analysis

Based on preliminary findings, the following cases are prioritized for Phase 2:

CaseReason for Selection
NG_001Type II Logic Failure candidate (5/6 correct but wrong reasoning – preliminary)
AU_013Universal failure – understand why all models wrong
UK_013Mixed performance – understand proprietary estoppel reasoning
US_013Most difficult recent case – understand recency effect
NG_009Procedural blindness – understand why models missed limitation point
NG_020Document hierarchy failure – understand why all models defaulted to wrong intuition

SECTION 18: NEW RESEARCH QUESTIONS EMERGING FROM THE FINDINGS

The findings presented in this report, while comprehensive and fully addressing the original research questions, open numerous avenues for further investigation. This section catalogs 35 new research questions that emerge directly from the data, organized by thematic area. These questions represent opportunities for future research, potential PhD extensions, and practical investigations that could further enhance understanding of LLM legal reasoning.

18.1 Model Behavior & Performance Questions

  • Why does Grok consistently decline with citation clues? Grok is the only model that performs worse with clues in multiple jurisdictions (UK: -2, Australia: -3, Nigeria: 0). Is this because its training data is structured differently such that citation retrieval introduces noise rather than signal? Does Grok over-rely on parametric knowledge and fail to integrate new contextual information?
  • What explains DeepSeek's extreme UK specialization? DeepSeek achieves perfect 20/20 on UK without clues but scores only 9/20 on US and Australia. Is its training data disproportionately UK-focused? Would it perform well on other Commonwealth jurisdictions (e.g., Canada, New Zealand, India)?
  • Why does Claude regress with clues in the US but improve elsewhere? Claude improves dramatically in UK (+4) and Australia (+5) but regresses in US (-3). Is its US legal knowledge stored differently such that citation retrieval triggers incorrect associations? Does its training data contain systematic errors in US Supreme Court jurisprudence?
  • What drives Perplexity's dramatic clue effect? Perplexity improves from 12/20 to 19/20 in UK (+7) but shows no improvement in Australia (0). Is its training data UK-heavy but Australia-light? Does it rely more heavily on retrieval than reasoning, and at what point does this retrieval dependency become a liability?
  • Are there "families" of models with similar behavior patterns? Grok and DeepSeek both decline with clues; ChatGPT and Claude both improve consistently. Do models from the same architectural "family" behave similarly? Can we cluster models by performance patterns and infer architectural features from these patterns?

18.2 Domain-Specific Questions

  • Why do procedural/limitation cases cause universal failure? NG_009 (statute-bar) achieved 0/6 in No-Clue and only 2/6 in With-Clue. Do LLMs systematically underweight procedural rules in favor of substantive merits? Is this a training data issue (procedural cases underrepresented) or an architectural issue (models biased toward substantive reasoning)?
  • Why does document hierarchy (parol evidence) cause universal failure? NG_020 (Atiba v Suberu) achieved 0/6 in No-Clue. Do LLMs lack understanding of integration clauses and document hierarchy? Is this because training data contains more examples of subsequent conduct modifying agreements than formal deed supremacy?
  • Why does purposive construction cause universal failure? AU_013 (Ecosse Property Holdings) achieved 0/6 in No-Clue, with all models applying literal textual construction. Do LLMs have a systematic bias toward literal interpretation over purposive construction? Can models be prompted to consider commercial purpose?
  • Which legal domains are most and least reliably handled? Domain performance varies dramatically (Contract interpretation 63.9% in Australia, Commercial law 88.9% in Australia). Is there a hierarchy of legal domains by LLM reliability? Which domains are most resistant to improvement with citation clues?
  • How do models perform on constitutional vs private law questions? This dataset focuses on private law. Do models perform differently on constitutional questions? Is there a difference between statutory interpretation and common law reasoning?
  • Do models perform better on landmark vs obscure cases? Some cases were universally correct, others universally wrong. Is there a correlation between case citation frequency in training data and model accuracy? Do models perform better on cases that appear frequently in legal education materials?

18.3 Temporal Questions

  • What is the precise training data cutoff for each model? The sharp performance decline for 2025 cases (61.5% vs 78.3% for 2015 cases) suggests recency effects. Can we estimate each model's training data cutoff date from performance trajectories? Do models have different cutoff dates?
  • How does case age interact with jurisdiction? Is the recency effect uniform across jurisdictions or more pronounced in some? Do models perform better on older cases from their "specialized" jurisdictions?
  • How will these findings change as models update? Will new model versions show improved performance on previously difficult cases? Does the recency effect persist across model generations? How quickly do models incorporate new jurisprudence?

18.4 Calibration & Confidence Questions

  • Why is DeepSeek's calibration perfect in UK but poor elsewhere? DeepSeek's NC Brier Score in UK is 0.0147 (excellent) but degrades in other jurisdictions. Is calibration jurisdiction-specific? Does poor calibration in unfamiliar jurisdictions signal lack of genuine understanding?
  • What causes the "overconfidence trap" in some models? DeepSeek expressed 100% confidence on multiple wrong answers (UK_004, UK_015), generating BS of 1.0. Which models are most prone to overconfidence? Does overconfidence correlate with training data gaps?
  • Do models know when they don't know? Some models (e.g., Grok) occasionally predicted incorrectly with lower confidence, others (DeepSeek) with high confidence. Which models have the best "metacognitive awareness"? Does calibration improve with citation clues? Can confidence scores be used to triage cases requiring manual review?

18.5 Methodological Questions

  • How much did the Translation Framework affect rankings? Without the Translation Framework, three Nigerian scores would have been incorrect. How many scores across all jurisdictions would be affected? Do some models systematically express predictions in substantive terms more than others?
  • How often does the Hard Boundary Rule apply? The rule was triggered once in this dataset. Are internally inconsistent responses rare or common? Do certain models produce more inconsistent responses? Should inconsistent responses be treated differently in scoring?
  • What is the optimal fact pattern length? Sanitized fact patterns were 75-90 words. Is there an optimal length for legal reasoning? Does information density matter more than absolute length? How much can facts be compressed without losing essential legal content?

18.6 Practitioner-Focused Questions

  • Can we create a jurisdiction-specific reliability index? Models have clear jurisdictional strengths and weaknesses. Can we create a "reliability score" for each model in each jurisdiction? How should practitioners weight these scores when choosing models? Should the index be updated as models evolve?
  • What is the cost-benefit of using multiple models? Different models excel in different areas. Does using an ensemble of models improve accuracy? What is the optimal ensemble strategy for legal research? How should practitioners reconcile conflicting predictions from different models?
  • Can we predict case difficulty for LLMs? Some cases were universally correct, others universally wrong. Can case characteristics (domain, year, procedural posture) predict LLM performance? Is there a "difficulty index" that correlates across models? Can practitioners pre-screen cases for AI reliability based on case metadata?
  • How should the Duty of Inquiry Checklist be customized by jurisdiction? Different jurisdictions pose different challenges (e.g., procedural blindness in Nigeria, purposive construction in Australia, document hierarchy in Nigeria). Should verification protocols be jurisdiction-specific? What are the top 3 verification priorities for each jurisdiction?

18.7 Comparative Questions

  • How do these models compare to human lawyers? Models achieve 60-100% accuracy depending on condition and jurisdiction. How would experienced lawyers perform on the same task? At what point does AI accuracy exceed typical human performance? Which cases are humans better at than AI?
  • How do these models compare to each other statistically? Descriptive differences are clear, but statistical significance needs testing. Which model differences are significant after correction for multiple comparisons? Is the ranking stable across different statistical tests? What sample size is needed to detect meaningful differences?
  • Are there "families" of models with similar behavior? Grok and DeepSeek both decline with clues; ChatGPT and Claude both improve. Do models from the same "family" (e.g., both retrieval-augmented) behave similarly? Can we cluster models by performance patterns? Does model architecture predict performance patterns?

18.8 Longitudinal Questions

  • How will these findings change as models update? Current snapshot as of early 2026. Will new model versions show improved performance on previously difficult cases? Does the recency effect persist across model generations? How quickly do models incorporate new jurisprudence?
  • Can fine-tuning on Nigerian law improve performance? ChatGPT achieves perfect score with clues but Grok leads without clues. How much would fine-tuning on Nigerian case law improve performance? Which model is most amenable to fine-tuning? What is the optimal fine-tuning dataset size and composition?

18.9 Theoretical Questions

  • What does "understanding" mean in the context of LLM legal reasoning? Type II Logic Failures (correct verdict, wrong reasoning) show that correctness ≠ understanding. Can we develop a metric for "genuine understanding" that goes beyond outcome prediction? How often do models achieve correct outcomes through incorrect reasoning? What does this tell us about how LLMs represent legal knowledge?
  • Is jurisdictional bias a form of epistemic injustice? Some jurisdictions (Nigeria, Australia) are systematically disadvantaged in model performance. Does the performance gap constitute a form of epistemic injustice? Are developing legal systems systematically disadvantaged in AI training data? What are the ethical implications of jurisdictionally biased AI tools?
  • What does the clue effect reveal about model architecture? Clue effect varies dramatically by model (Gemini +13, Perplexity +9, Grok -3). Does the clue effect correlate with model architecture (e.g., retrieval-augmented vs pure parametric)? Can we infer architectural differences from performance patterns? How do different models integrate retrieved information with parametric knowledge?

18.10 Practical Implementation Questions

  • Can we automate the Duty of Inquiry Checklist? Manual verification is essential but time-consuming. Can we build a tool that automatically verifies citations? Can we flag potential Type II Logic Failures algorithmically? What is the optimal human-AI collaboration model for legal research?
  • How should law firms integrate these findings into practice? Clear guidance on model selection and input preparation. What is the best way to communicate findings to practicing lawyers? Should firms develop internal AI usage policies based on these results? How can the Duty of Inquiry Checklist be integrated into existing workflows?
  • What is the economic impact of using optimal vs suboptimal models? Model performance varies by 20-40% across jurisdictions. What is the cost of using the wrong model for a given jurisdiction? How many hours of lawyer time could be saved by optimal model selection? What is the ROI of implementing the Duty of Inquiry Checklist?

18.11 Summary: A Research Agenda

These 35 questions represent not merely extensions of this study, but an entire research program. They cluster into several natural research streams:

Research StreamQuestionsPotential Outputs
Model Architecture & Behavior1-5, 27, 32Journal articles on model design
Domain-Specific Legal Reasoning6-11Domain reliability indices
Temporal Dynamics12-14, 28-29Training cutoff estimation methods
Calibration & Confidence15-17Confidence-based reliability tools
Methodological Innovation18-20Improved benchmarking protocols
Practical Applications21-24, 33-35Practitioner tools and guidelines
Comparative Studies25-26Human-AI comparison studies
Theoretical Foundations30-31Epistemic justice in AI

Priority Recommendations for Immediate Follow-up:

  • Investigate Grok's clue-induced decline – This anomalous finding (-3 overall) could reveal important insights about model architecture and retrieval mechanisms.
  • Analyze DeepSeek's UK specialization – Understanding this could help predict model performance in other Commonwealth jurisdictions.
  • Study procedural blindness – The universal failure on NG_009 has immediate practical implications for legal practitioners.
  • Study document hierarchy failure – The universal failure on NG_020 reveals a critical weakness in Nigerian property and banking law contexts.
  • Develop jurisdiction-specific reliability indices – Directly useful for practitioners and could become a standard tool.
  • Conduct human lawyer comparison study – Essential for contextualizing findings and establishing practical relevance.

SECTION 19: CONCLUSION

19.1 What This Research Has Achieved

AchievementEvidence
Scale960 verified experiments across 4 jurisdictions
RigorFull audit trail, transparent methodology, documented corrections
OriginalityFirst cross-jurisdictional LLM legal study including Nigeria
Practical impactFirst evidence-based AI guidance for Nigerian legal profession
Theoretical contributionPreliminary empirical support for "stochastic parrot" hypothesis (awaiting Phase 2 confirmation); identification of "citation interference" pattern
Methodological innovationTranslation Framework, Hard Boundary Rule, Two-Step Verification
Domain mappingFirst detailed performance mapping by legal domain
Recency quantification16.8% accuracy drop for 2025 cases documented
Failure pattern identificationProcedural blindness, document hierarchy confusion, purposive construction weaknesses, citation interference
Industry Translation: Eight deployable governance outputsFirst operational framework for Nigerian legal AI adoption in a regulatory vacuum.

Industrial Impact Summary:

This research produces nine deployable outputs for the Nigerian legal sector: (1) Traceable Accountability three-pillar framework (Section 14.5); (2) risk-classified model selection matrix with explicit warnings (Section 14.1); (3) Nigerian legal domain risk map (Appendix G); (4) five-phase implementation roadmap (Section 14.6); (5) vendor selection scorecard (Appendix D); (6) expanded Duty of Inquiry Verification Checklist (Section 14.4); (7) AI Governance Committee charter template (Appendix E); (8) incident tracking protocol (Appendix F). These outputs translate 960 empirical observations into actionable IT management tools, enabling Nigerian law firms to deploy LLMs with documented accountability despite the absence of binding AI regulation.

The eight purely empirical outputs are:

#OutputEmpirical Source
1Traceable Accountability Three Pillar FrameworkRisk patterns from your data
2Risk Classified Model Selection MatrixAccuracy tables (Sections 3–7) + statistical validation
3Nigerian Legal Domain Risk MapSection 10.1 domain accuracy data
4Five Phase Implementation RoadmapLogical sequence from your risk management framework (Section 14.3) – no external statistics
5Vendor Selection ScorecardWeighted criteria from your metrics (Brier, EV, CV, procedural recovery)
6Expanded Duty of Inquiry ChecklistUniversal failure cases (NG_009, NG_020) + Type II Logic Failure
7AI Governance Committee Charter TemplateDerived from your risk management framework
8Incident Tracking and Remediation ProtocolDerived from error types you observed (hallucinations, procedural blindness, overconfidence)

19.2 Key Findings Summary (Descriptive) – FINAL CORRECTED

  • Grok is most jurisdiction-independent without clues (83.8% across 4 jurisdictions)
  • Gemini is most reliable with citation context (93.8% across 4 jurisdictions)
  • DeepSeek shows extreme UK specialization (20/20 UK vs 9/20 US) but ⚠️ regresses with clues and exhibits overconfidence on errors
  • ChatGPT achieves perfect score on Nigerian law with clues (20/20)
  • Perplexity shows largest clue effect (+7 in UK) but remains retrieval-dependent
  • Type II Logic Failure (preliminary) – models may be correct for wrong reasons; Phase 2 verification required
  • Universal failures reveal systematic gaps in:
  • Procedural/limitation law (NG_009)
  • Document hierarchy/parol evidence (NG_020)
  • Purposive construction (AU_013)
  • Citation clues help most models (+9.1% avg) but harm Grok (-3 overall)
  • Grok is the only model to consistently decline with clues across multiple jurisdictions (UK -2, Australia -3, Nigeria 0)
  • Claude shows largest improvement in Australia (+5) and UK (+4)
  • Recency effect – 16.8% accuracy drop for 2025 cases
  • Domain-specific weaknesses – procedural law, document hierarchy, purposive construction, equity most challenging

19.3 Next Steps Timeline

PhaseActivityEstimated Duration
1Statistical testing (Friedman, Wilcoxon, McNemar, etc.)10-12 days
2Researcher analysis of Key Reason (Phase 1)2 weeks
3Lawyer verification of Key Reason (Phase 2)1 week (parallel)
4Type II Logic Failure quantification3 days
5Practitioner Manual drafting1 week
6Thesis chapter writing (Results & Discussion)2 weeks
7Industrial toolkit finalisation | Practitioner Manual + 9 outputs1 week (parallel to Phase 2)

19.4 Dual Contribution Statement

This research makes two distinct but interconnected contributions:

Academic Contribution: First systematic cross-jurisdictional LLM legal reasoning benchmark including an African common-law jurisdiction (Nigeria) alongside the US, UK, and Australia. Introduces and operationalises Type II Logic Failure (correct verdict, wrong reasoning) as a metric distinguishing genuine understanding from stochastic parroting. Provides statistical confirmation of model-specific jurisdiction sensitivity, citation interference, and recency effects.

Industrial Contribution: First operational governance framework and practitioner toolkit for responsible AI adoption in a regulatory-vacuum jurisdiction. Delivers nine deployable outputs (enumerated above) that translate empirical findings into IT management tools, procurement criteria, and professional verification standards. Provides evidence-based guidance to the Nigerian Bar Association, NITDA, and Nigerian Law Schools for policy development.

The two contributions are mutually reinforcing: the academic benchmark provides the empirical foundation; the industrial toolkit provides the implementation pathway. Neither is complete without the other.

Future Research Direction – Operationalising Relational AI Governance:

While this study provides an empirically grounded risk management framework, the literature suggests that African ethical values such as Ubuntu (relationality, communal well being) could inform AI governance in culturally resonant ways (Mutswiri et al., 2025; Yilma, 2025). However, Ubuntu’s normative structure remains under specified for direct implementation (Mensah & Van Wynsberghe, 2025). Future research should translate such principles into concrete institutional rules, building on the Traceable Accountability framework developed here. The current study does not claim Ubuntu as a validated contribution.

Report Appendices A–G · Master sheet, sample rows, correction log, governance templates, Nigerian domain risk mapReport lines 1109–1316

APPENDIX A: COMPLETE MASTER SHEET STRUCTURE

ColumnHeaderDescription
AExperiment_IDUnique identifier: CaseID_Condition_Model
BCase_IDOriginal case identifier
CJurisdictionCountry
DConditionNo Clue or With Clue
EModelLLM name
FVerdict_Score1 = correct, 0 = incorrect
GConfidenceConfidence percentage (0–100)
HBrier_ScoreWhen correct: (Confidence/100 - 1)²

When incorrect: (Confidence/100 - 0)²

IExpected_Value+Confidence/100 if correct, –Confidence/100 if wrong
JCase_TitleFull citation
KLegal_DomainArea of law
LYearJudgment year
MGround_Truth_ReasoningResearcher-extracted ratio decidendi (with references)
NReasoning_Classification_PrelimResearcher's preliminary classification
OReasoning_Classification_FinalLawyer's verified classification
PReasoning_TypeAligned / General-but-Correct / General-but-Wrong / Misaligned
QType_II_Logic_FailureYes/No flag for General-but-Correct with jurisdiction-specific principle
RLawyer_NotesAdditional comments from legal practitioner
SLawyer_Verified_DateDate of verification

APPENDIX B: SAMPLE MASTER SHEET ROWS

Experiment_IDCase_IDJurisdictionConditionModelVerdict_ScoreConfidenceBrier_ScoreExpected_ValueCase_TitleLegal_DomainYear
NG_001_NoClue_ChatGPTNG_001NigeriaNo ClueChatGPT1780.04840.78Damisa v. U.B.A. (2025)Contract Law2025
NG_001_NoClue_GeminiNG_001NigeriaNo ClueGemini1950.00250.95Damisa v. U.B.A. (2025)Contract Law2025
NG_001_NoClue_ClaudeNG_001NigeriaNo ClueClaude1720.07840.72Damisa v. U.B.A. (2025)Contract Law2025
NG_005_WithClue_DeepSeekNG_005NigeriaWith ClueDeepSeek01001.0000-1.00Damisa v. U.B.A. (2025)Contract Law2025
AU_018_WithClue_GrokAU_018AustraliaWith ClueGrok0750.5625-0.75Westpac v. Linside (2020)Commercial Law2020

APPENDIX C: CORRECTION LOG

DateSectionOriginal ValueCorrected ValueReason
March 10, 20266.2 Australia With Clue – Grok total1312AU_018 Grok score corrected from 1 → 0 (post-hoc verification)
March 10, 20266.3 Australia With Clue Accuracy – Grok65.0%60.0%Cascade from above
March 10, 20267.2 Cross-jurisdictional With Clue – Grok total6564Cascade from above
March 10, 20267.2 Cross-jurisdictional With Clue – Grok avg81.3%80.0%Cascade from above
March 10, 20267.3 Clue Effect – Grok improvement-2-3Cascade from above
March 10, 20267.3 Mean improvement all models+9.4%+9.1%Cascade from above

END OF COMPLETE RESEARCH FINDINGS REPORT – FINAL CORRECTED VERSION

All 960 observations verified and audited. Single inconsistency identified and corrected. All totals internally consistent. Statistical testing ready. Phase 2 (Reasoning Analysis) prepared. Practitioner Manual draftable.

APPENDIX D: VENDOR SELECTION SCORECARD

Use this scorecard to evaluate LLM vendors for Nigerian legal practice. Minimum passing score: 75/100.

CriterionWeightScoring GuideChatGPTGeminiClaudeGrokDeepSeekPerplexity
Nigerian accuracy with citations30%100%=30; 95%=28.5; 90%=27; 85%=25.5; 80%=24; <80%=030 (100%)28.5 (95%)25.5 (85%)24 (80%)19.5 (65%)24 (80%)
Nigerian accuracy without citations20%80%=16; 75%=15; 70%=14; 65%=13; 60%=12; <60%=014 (70%)12 (60%)14 (70%)16 (80%)10 (50%)15 (75%)
Cross-jurisdictional consistency15%CV <0.10=15; 0.10-0.15=12; 0.15-0.20=9; >0.20=6; >0.40=012 (0.1242)6 (0.2140)9 (0.1695)15 (0.0896)0 (0.4462)12 (0.1091)
High-confidence error rate15%<5%=15; 5-10%=12; 10-15%=9; 15-20%=6; >20%=015 (<5%)15 (<5%)12 (~8%)12 (~8%)0 (>20%)12 (~8%)
Procedural/domain recovery10%Corrected both=10; one=5; none=010 (both)10 (both)0 (neither)5 (one)0 (neither)5 (one)
Transparency/disclosure10%Full=10; partial=5; none=0555505
TOTAL SCORE100%8676.565.57729.573
RecommendationAPPROVEAPPROVECONDITIONALAPPROVEREJECTCONDITIONAL

Minimum Passing Score: 75/100

Approved Vendors (≥75): ChatGPT (86), Grok (77), Gemini (76.5)

Conditional Approval (65-74): Perplexity (73), Claude (65.5) – requires enhanced verification protocol

Rejected (<65): DeepSeek (29.5) – negative EV, high-confidence error rate, no procedural recovery

Contractual Clauses for Approved Vendors:

1. Vendor must provide quarterly accuracy reports on held-out Nigerian cases

2. Material performance degradation (>15% drop) constitutes breach

3. Vendor must disclose training data updates and recency cutoffs

APPENDIX E: AI GOVERNANCE COMMITTEE CHARTER TEMPLATE

[LAW FIRM NAME] AI GOVERNANCE COMMITTEE CHARTER

Effective Date: _______________

1. Purpose

The AI Governance Committee oversees the responsible adoption, deployment, and monitoring of artificial intelligence tools in legal practice, ensuring compliance with professional obligations and risk management standards derived from empirical benchmarking (Uba, 2026).

2. Membership

  • Senior Partner (Chair) – one position
  • IT Manager – one position
  • Training Officer – one position
  • External Ethics Advisor – one position (rotating quarterly)
  • Legal Practitioner Representative – one position (rotating annually)

3. Responsibilities

  • Evaluate LLM vendors using the Vendor Selection Scorecard (Appendix D)
  • Maintain and update the Restricted Model List (models requiring case-by-case approval)
  • Review and approve AI Use Policy amendments
  • Conduct quarterly tool reviews (accuracy, incident rates, new model versions)
  • Review incident reports (see Appendix F) and recommend remediation
  • Ensure annual AI literacy training is delivered
  • Liaise with NBA and NITDA on policy developments

4. Meeting Schedule

  • Quarterly: Full committee review (2 hours)
  • Monthly: Subcommittee on incidents (1 hour, as needed)

5. Reporting

  • Quarterly report to firm partners (template available)
  • Annual governance audit report
  • Incident summaries to training officer for CLE materials

6. Decision Authority

  • Vendor approval/rejection (requires 2/3 majority)
  • Restricted list additions (requires Chair approval)
  • Emergency suspension of any AI tool (IT Manager + Chair)

Signatures:

_________________ (Chair) Date: _______________

_________________ (IT Manager) Date: _______________

APPENDIX F: INCIDENT TRACKING TEMPLATE

AI-Related Incident Report Form

| Field | Entry |

|-------|-------|

| Incident ID | AI-[YYYY]-[XXX] |

| Date of incident | _______________ |

| Reporting lawyer | _______________ |

| Model used | ChatGPT / Gemini / Claude / Grok / DeepSeek / Perplexity / Other: ______ |

| Condition at time | With citations / Without citations / Partial clue |

| Case jurisdiction | Nigeria / UK / US / Australia / Other: ______ |

| Legal domain | Contract / Tort / Commercial / Procedural / Property / Equity / Other: ______ |

| Case year | _______________ |

Description of Incident

[What was the AI output? What was the error?]

Error Type (select all that apply)

[ ] Hallucinated citation (non-existent case)

[ ] Wrong jurisdiction applied

[ ] Procedural blindness (missed limitation period)

[ ] Document hierarchy error (parol evidence)

[ ] Purposive construction error (literal interpretation)

[ ] Overconfidence (high confidence on wrong answer)

[ ] Type II Logic Failure (correct verdict, wrong reasoning)

[ ] Other: _______________

Client Impact

[ ] No client impact (caught internally)

[ ] Client advised incorrectly – corrected before filing

[ ] Filed with error – corrected after filing

[ ] Adverse outcome for client

Remediation Actions

[ ] Manual verification performed

[ ] Client notified

[ ] Court notified (if applicable)

[ ] Training updated

[ ] Model added to restricted list (temporary/permanent)

Submitted to AI Governance Committee on: _______________

Committee Resolution: _______________

Follow-up actions: _______________

APPENDIX G: NIGERIAN LEGAL DOMAIN RISK MAP (COLOUR-CODED)

🟢 GREEN – LOW RISK (Safe for AI delegation with routine verification)

Focus: Areas where LLMs demonstrate near-perfect alignment with Nigerian jurisprudence.

DomainAccuracy (NC)Verification Required
Criminal Law100%Standard outcome check against official reports.

🟡 YELLOW – MEDIUM RISK (AI use acceptable with standard verification including reasoning)

Focus: General common law principles that transfer well but require "Reasoning Audits."

DomainAccuracy (NC)Verification Required
Tort Law72.0%Cross-verify outcome + legal reasoning logic.
Contract Law (Labour)77.8%Check for Type II Logic Failure (Right answer, wrong reason).
Contract Law (Non-land)75% (est.)Verify specific jurisdiction-dependent principles.

🟠 ORANGE – HIGH RISK (AI use only with mandatory manual verification)

Focus: Highly specialized Nigerian doctrines where models frequently hallucinate or apply US/UK rules.

DomainAccuracy (NC)Verification Required
Contract Law (Land/Mortgage)58.3%Mandatory manual verification of document hierarchy.
Commercial Law (Banking/IP)66.7%Mandatory procedural check for limitation periods.
Property/Equity58.3%Manual verification of Nigeria-specific equitable rules.

🔴 RED – CRITICAL RISK (AI should NOT be used without independent verification)

Focus: Structural failure zones where LLM architecture is "Procedurally Blind."

DomainAccuracy (NC)Verification Required
Procedural/Limitation Law0% (NG_009)DO NOT RELY ON AI. Manual verification is mandatory.
Document Hierarchy/Parol Evid.0% (NG_020)DO NOT RELY ON AI. Manual verification is mandatory.

One-page Quick Reference for Practitioners

ColourMeaningAction to be Taken
🟢 GREENSafe to delegateUse AI; verify the final outcome.
🟡 YELLOWUse with cautionVerify the outcome AND the underlying reasoning.
🟠 ORANGEHigh riskMandatory manual verification of specific technical elements.
🔴 REDCritical riskDO NOT rely on AI. Perform manual research only.

This risk map is derived from Section 10.1 (Nigeria) and universal failure cases (Section 13). Update annually based on new model versions.

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.

Benchmarking report · where everything sits

Map of the full benchmarking report

Every empirical section of the report is placed on the study page it informs. The only text not reproduced is an argumentative essay on the study’s long-term relevance and a research-ethics reflection, which contain no results.

  1. Executive summary, headline findings and timeline · appendices
  2. Section 1–2 · Methodology, data integrity and research framework · benchmark method
  3. Sections 3–6 · Complete score sheets for Nigeria, UK, US and Australia · benchmark accuracy
  4. Section 7 · Cross-jurisdictional comparison and clue effect by model · citation context
  5. Sections 8–9 · Model specialisation by jurisdiction and Commonwealth lineage · benchmark accuracy
  6. Section 10 · Performance by legal domain · benchmark accuracy
  7. Section 11 · Performance by case year · decision year
  8. Section 12 · The “correct guessing” observation · preliminary reasoning
  9. Section 13 · Notable failure cases · confidence
  10. Section 14 · Recommendations and governance tools for practitioners (planned resources) · next steps
  11. Sections 15–19 · Theoretical contributions, statistical plan, Phase 2, emerging questions and conclusion · appendices
  12. Report Appendices A–G · Master sheet, sample rows, correction log, governance templates, Nigerian domain risk map · appendices
  13. Part 2 · Statistical confirmation of the benchmark (report-only tests) · citation context
  14. Part 3 · From statistical confirmation to follow-up experiment design · continuation experiments
  15. Continuation experiments report · Experiments 1–3 in full · continuation experiments
  16. Phase 2 · Reasoning analysis (researcher key-reason review, awaiting lawyer verification) · preliminary reasoning
  17. Report Appendix A · Supplementary statistical analyses (logistic regression, ECE, domain chi-square, power) · confidence
  18. Silent Failure analysis and findings report · confidence

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 12.1 · source-reported

Mapping between the labels used in this thesis and the original labels

      Label             Words/source    Original label    Stage

      GLMRS-A           50–100          GLMRS-D (prov.)   Third
      GLMRS-B           100–150         GLMRS-E (prov.)   Third
      GLMRS-C           150–300         GLMRS-A           Second
      GLMRS-D           300–500         GLMRS-B           Second
      GLMRS-E           500–750         GLMRS-C           Second
      GLMRS No Clue     150–300         Unchanged         First (simultaneous)

Thesis Table 12.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.2 · source-reported

Fourth stage, run a, 100–150 words per source: ChatGPT verdicts by summarisation tool

    Case      GT       ChatGPT   Claude    Copilot   Gemini Gemini Notebook Correct /5

    NG 001    Dism.    D 58% ✓   D 68% ✓   D 70% ✓   D 62% ✓    D 60% ✓          5
    NG 002    Dism.    D 51% ✓   D 52% ✓   A 52% ×   A 51% ×    A 52% ×          2
    NG 003    Dism.    D 74% ✓   D 76% ✓   D 78% ✓   D 74% ✓    D 72% ✓          5
    NG 004    Dism.    D 52% ✓   D 50% ✓   D 51% ✓   D 53% ✓    D 52% ✓          5
    NG 005    Dism.    D 57% ✓   D 58% ✓   D 66% ✓   D 57% ✓    D 57% ✓          5
    NG 006    Allow.   A 53% ✓   A 50% ✓   A 51% ✓   A 50% ✓    A 51% ✓          5
    NG 007    Dism.    D 62% ✓   D 72% ✓   D 76% ✓   D 58% ✓    D 58% ✓          5
    NG 008    Dism.    D 67% ✓   D 58% ✓   D 73% ✓   D 69% ✓    D 68% ✓          5
    NG 009    Dism.    A 61% ×   A 56% ×   A 67% ×   A 67% ×    A 64% ×          0
    NG 010    Dism.    D 58% ✓   D 62% ✓   D 61% ✓   D 59% ✓    D 60% ✓          5
    NG 011    Allow.   A 82% ✓   A 84% ✓   A 88% ✓   D 88% ×    A 78% ✓          4
    NG 012    Dism.    D 50% ✓   D 50% ✓   D 50% ✓   D 51% ✓    D 51% ✓          5
    NG 013    Allow.   A 51% ✓   A 52% ✓   A 52% ✓   A 52% ✓    A 53% ✓          5
    NG 014    Allow.   A 54% ✓   A 57% ✓   A 57% ✓   A 54% ✓    A 55% ✓          5
    NG 015    Dism.    D 52% ✓   D 53% ✓   D 54% ✓   D 53% ✓    D 52% ✓          5
    NG 016    Dism.    D 88% ✓   D 83% ✓   D 86% ✓   D 86% ✓    D 82% ✓          5
    NG 017    Allow.   A 52% ✓   A 51% ✓   D 50% ×   A 50% ✓    A 50% ✓          4
    NG 018    Dism.    D 72% ✓   D 66% ✓   A 82% ×   D 70% ✓    D 67% ✓          4
    NG 019    Allow.   A 52% ✓   A 50% ✓   A 58% ✓   D 54% ×    D 55% ×          3
    NG 020    Dism.    A 53% ×   A 61% ×   A 55% ×   A 63% ×    A 56% ×          0

    Correct             18/20     18/20     15/20     15/20      16/20        82/100

Thesis Table 12.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.3 · source-reported

Fourth stage, run a, 150–300 words per source: ChatGPT verdicts by summarisation tool

    Case      GT       ChatGPT   Claude    Copilot   Gemini Gemini Notebook Correct /5

    NG 001    Dism.    D 62% ✓   D 72% ✓   D 72% ✓   D 66% ✓    D 61% ✓          5
    NG 002    Dism.    D 52% ✓   D 52% ✓   A 53% ×   A 52% ×    A 53% ×          2
    NG 003    Dism.    D 82% ✓   D 81% ✓   D 82% ✓   D 78% ✓    D 78% ✓          5
    NG 004    Dism.    D 52% ✓   D 50% ✓   D 51% ✓   D 52% ✓    D 52% ✓          5
    NG 005    Dism.    D 61% ✓   D 64% ✓   D 69% ✓   D 59% ✓    D 59% ✓          5
    NG 006    Allow.   A 54% ✓   A 63% ✓   A 51% ✓   A 50% ✓    A 54% ✓          5
    NG 007    Dism.    D 66% ✓   D 71% ✓   D 79% ✓   D 61% ✓    D 59% ✓          5
    NG 008    Dism.    D 70% ✓   D 63% ✓   D 78% ✓   D 72% ✓    D 71% ✓          5
    NG 009    Dism.    A 67% ×   A 59% ×   A 72% ×   A 72% ×    A 67% ×          0
    NG 010    Dism.    D 63% ✓   D 71% ✓   D 64% ✓   D 62% ✓    D 63% ✓          5
    NG 011    Allow.   D 91% ×   D 89% ×   A 90% ✓   D 92% ×    A 75% ✓          2
    NG 012    Dism.    D 50% ✓   D 50% ✓   D 50% ✓   D 52% ✓    D 51% ✓          5
    NG 013    Allow.   A 51% ✓   A 54% ✓   A 54% ✓   A 57% ✓    A 54% ✓          5
    NG 014    Allow.   A 55% ✓   A 60% ✓   A 58% ✓   A 56% ✓    A 57% ✓          5
    NG 015    Dism.    D 53% ✓   D 54% ✓   D 55% ✓   D 55% ✓    D 53% ✓          5
    NG 016    Dism.    D 93% ✓   D 88% ✓   D 89% ✓   D 91% ✓    D 85% ✓          5
    NG 017    Allow.   A 52% ✓   A 53% ✓   D 50% ×   A 51% ✓    A 50% ✓          4
    NG 018    Dism.    D 82% ✓   D 72% ✓   A 86% ×   D 76% ✓    D 72% ✓          4
    NG 019    Allow.   A 53% ✓   A 51% ✓   A 59% ✓   D 56% ×    D 56% ×          3
    NG 020    Dism.    A 55% ×   A 67% ×   A 60% ×   A 68% ×    A 61% ×          0

    Correct             17/20     17/20     15/20     15/20      16/20        80/100

Thesis Table 12.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.4 · source-reported

Fourth stage, run b (reruns): ChatGPT verdicts with ChatGPT and Claude summaries

                                               100–150 words           150–300 words

                      Case       GT           ChatGPT     Claude      ChatGPT      Claude

                      NG 001     Dism.        D 64% ✓     D 62% ✓     D 66% ✓      D 65% ✓
                      NG 002     Dism.        A 52% ×     D 35% ✓     A 52% ×      D 35% ✓
                      NG 003     Dism.        D 76% ✓     D 73% ✓     D 82% ✓      D 78% ✓
                      NG 004     Dism.        D 53% ✓     D 36% ✓     D 53% ✓      D 36% ✓
                      NG 005     Dism.        D 67% ✓     D 46% ✓     D 70% ✓      A 57% ×
                      NG 006     Allow.       A 55% ✓     A 28% ✓     A 56% ✓      A 30% ✓
                      NG 007     Dism.        D 79% ✓     D 55% ✓     D 82% ✓      D 64% ✓
                      NG 008     Dism.        D 72% ✓     D 44% ✓     D 77% ✓      D 53% ✓
                      NG 009     Dism.        A 70% ×     D 43% ✓     A 75% ×      D 52% ✓
                      NG 010     Dism.        D 61% ✓     D 48% ✓     D 65% ✓      D 54% ✓
                      NG 011     Allow.       A 86% ✓     A 78% ✓     D 92% ×      D 72% ×
                      NG 012     Dism.        D 51% ✓     D 38% ✓     D 51% ✓      D 38% ✓
                      NG 013     Allow.       A 54% ✓     A 39% ✓     A 56% ✓      A 40% ✓
                      NG 014     Allow.       A 58% ✓     A 41% ✓     A 59% ✓      A 42% ✓
                      NG 015     Dism.        D 55% ✓     D 39% ✓     D 55% ✓      D 39% ✓
                      NG 016     Dism.        D 88% ✓     D 82% ✓     D 92% ✓      D 88% ✓
                      NG 017     Allow.       A 50% ✓     A 31% ✓     A 50% ✓      A 32% ✓
                      NG 018     Dism.        D 66% ✓     D 58% ✓     D 72% ✓      D 67% ✓
                      NG 019     Allow.       A 58% ✓     A 32% ✓     A 61% ✓      A 33% ✓
                      NG 020     Dism.        A 62% ×     A 40% ×     A 70% ×      A 48% ×

                      Correct                  17/20          19/20    16/20        17/20

 Confidence values below 50% in the Claude-summary reruns are inconsistent with a stated
                          binary verdict (see the scoring notes).

Thesis Table 12.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.5 · source-reported

Word count of each stored summary (summary text only, exclud- ing the title)

                                                 100–150 words                      150–300 words

    No. Source                          GPT Claude Copil. Gemini GNB GPT Claude Copil. Gemini GNB

       1 Constitution 1999              127     150     117     144    144   213   293      185   255    231
       2 Labour Act                     136     150     136     145    144   216   293      200   253    200
       3 Trade Marks Act                146     149     138     144    144   238   293      209   274    181
       4 Land Use Act                   141     147     145     141    137   230   296      205   266    185
       5 BOFIA 2020                     141     146     139     136   n/a∗   207   299      191   267   n/a∗
       6 CBN Cons. Prot. Framework      134     150     136     143    127   216   287      194   277    174
       7 Admiralty Jurisdiction Act     141     145     137     138    135   246   295      200   282    155
       8 Sheriffs & Civil Process Act   139     150     132     140    148   228   297      192   289    165
       9 NAFDAC Act                     140     149     128     139    142   219   295      173   281    184
      10 Counterfeit & Fake Drugs Act   134     148     134     139    147   226   299      164   286    181
      11 Merchant Shipping Act          145     148     131     146    132   222   298      170   288    171

GPT = ChatGPT, Copil. = Microsoft Copilot, GNB = Gemini Notebook. ∗ Gemini Notebook
 was supplied with BOFIA but did not produce a usable BOFIA summary at either length.

Thesis Table 12.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.6 · source-reported

Benchmark cases and Ground Truth

              Case      Short title                      Ground Truth

              NG 001    Damisa v. U.B.A.                 Appeal Dismissed
              NG 002    Super Ceramics v. H.E.P. Eng.    Appeal Dismissed
              NG 003    Moore Associates v. Exphar       Appeal Dismissed
              NG 004    F.H.A. v. Oyedeji                Appeal Dismissed
              NG 005    N.Y.S.C. v. Ukachukwu            Appeal Dismissed
              NG 006    Akaolisa v. Okuma                Appeal Allowed
              NG 007    Skye Bank v. Adegun              Appeal Dismissed
              NG 008    Heritage Bank v. Bentworth       Appeal Dismissed
              NG 009    Ethiopian Airlines v. Polaris    Appeal Dismissed
              NG 010    Austin Laz v. GTBank             Appeal Dismissed
              NG 011    Glenyork v. Panalpina            Appeal Allowed
              NG 012    Olaniran v. Adebayo              Appeal Dismissed
              NG 013    ACMEL v. First Bank              Appeal Allowed
              NG 014    Total E&P v. Okwu                Appeal Allowed
              NG 015    A.B.C. Transport v. Omotoye      Appeal Dismissed
              NG 016    Barewa Pharm. v. F.R.N.          Appeal Dismissed
              NG 017    Standard Chartered v. Ameh       Appeal Allowed
              NG 018    BPS Eng. v. F.R.M.A.             Appeal Dismissed
              NG 019    Omni Products v. Union Bank      Appeal Allowed
              NG 020    Atiba Iyalamu v. Suberu          Appeal Dismissed
                   Fourteen cases were dismissed and six were allowed.

Thesis Table 12.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 12.7 · source-reported

Project resources

Resource              Specification                              Cost        Status
                                                                 (AUD)

LLM access            Six models, 960 primary interactions and   About 300   Expended
                      the continuation experiments
NWLR Online access    Premium Nigerian law reports               In kind     Secured
Lawyer verification   Two practising lawyers for the Phase 4     In kind     Secured
                      reasoning review
Contingency           Additional model access and verification   100         Proposed
                      support if required

Thesis Table 12.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.