Key takeaways
- Appendix tables retain their thesis numbering.
- Table 12.6 lists benchmark cases and ground truth.
- The earlier reports’ 35 open questions are reproduced separately from the thesis’s 11 RQs.
What was tested & why
Appendices
Thesis tables appear below. The newly supplied historical reports also contain a 35-question follow-up catalogue, reproduced separately here as report-only material. Their per-case Part I score sheets, preliminary P/F/N/A reasoning grades and supplementary outputs remain outside the verified thesis evidence set.
Report-only questions are not answers or established findings; reasoning grades await lawyer verification.
Source: Thesis Chapter 12; historical benchmarking reports · source-reported unless otherwise noted.
Historical benchmarking and weekly reports · report-only
35 open questions, not 35 thesis RQs
The earlier reports include this catalogue of possible follow-up questions. Some prompts presuppose explanations not established by the final thesis, and items 5/27 and 14/28 overlap. They are reproduced as questions, not answered findings, endorsed hypotheses, or the thesis’s 11 research questions.
Why does Grok consistently decline with citation clues? Is its training data structured differently such that citation retrieval introduces noise rather than signal?
What explains DeepSeek’s extreme UK specialization? Would it perform well on other Commonwealth jurisdictions (e.g., Canada, New Zealand, India)?
Why does Claude regress with clues in the US but improve elsewhere?
What drives Perplexity’s dramatic clue effect? Does it rely more heavily on retrieval than reasoning?
Are there ‘families’ of models with similar behavior patterns?
Why do procedural/limitation cases cause universal failure?
Why does document hierarchy (parol evidence) cause universal failure?
Why does purposive construction cause universal failure?
Which legal domains are most and least reliably handled?
How do models perform on constitutional vs. private law questions?
Do models perform better on landmark vs. obscure cases?
What is the precise training data cutoff for each model?
How does case age interact with jurisdiction?
How will these findings change as models update?
Why is DeepSeek’s calibration perfect in UK but poor elsewhere?
What causes the ‘overconfidence trap’ in some models?
Do models know when they don’t know? Which models have the best ‘metacognitive awareness’?
How much did the Translation Framework affect rankings?
How often does the Hard Boundary Rule apply?
What is the optimal fact pattern length?
Can we create a jurisdiction-specific reliability index?
What is the cost-benefit of using multiple models?
Can we predict case difficulty for LLMs?
How should the Duty of Inquiry Checklist be customized by jurisdiction?
How do these models compare to human lawyers?
How do these models compare to each other statistically?
Are there ‘families’ of models with similar behavior?
How will these findings change as models update?
Can fine-tuning on Nigerian law improve performance?
What does ‘understanding’ mean in the context of LLM legal reasoning?
Is jurisdictional bias a form of epistemic injustice?
What does the clue effect reveal about model architecture?
Can we automate the Duty of Inquiry Checklist?
How should law firms integrate these findings into practice?
What is the economic impact of using optimal vs. suboptimal models?
Weekly benchmarking report, “New Research Questions Emerging from the Findings,” pp. 40–41. Report-only; the thesis §1.3 and §10.9 supply the authoritative 11 RQs and answers.
Supplementary records held apart
The reports also contain per-case Part I score matrices, Nigeria With Clue researcher P/F/N/A reasoning grades, and supplementary statistical outputs. These have not been independently reconciled cell by cell against the final thesis. The P/F/N/A entries are not lawyer-verified reasoning-quality findings (thesis §3.12); no Type II Logic Failure rate is inferred from them.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Executive summary, headline findings and timelineReport lines 15–48
EXECUTIVE SUMMARY
This report presents the complete, verified findings from a comprehensive benchmarking study evaluating six leading large language models (ChatGPT, Gemini, Claude, Grok, DeepSeek, Perplexity) on their ability to predict appellate court outcomes across four common law jurisdictions: Nigeria, United Kingdom, United States, and Australia.
1. The Master Research Timeline (For your Presentation Slide)
This is the high-level overview perfect for Slide 10 of your presentation.
YEAR 1: FOUNDATION & EXPERIMENTATION | Phase | Activity | Status | | :--- | :--- | :--- | | Literature & design | Spring 2025 – Ongoing | ✅ Continuous (50+ papers reviewed & 20+ in-depth study) | | Case selection & extraction | Summer 2025/2026 – Autumn 2026 | ✅ Complete | | LLM testing (960 experiments) | Autumn 2026 | ✅ Complete | | Scoring & descriptive analysis | Autumn 2026 | ✅ Complete | | Statistical analysis (Friedman, McNemar, etc.) | Autumn 2026 | ✅ Complete | | Adaptive follow up experiments (addressing emerged patterns) | Autumn 2026 | Ongoing |
YEAR 2: VERIFICATION & TRANSLATION | Phase | Activity | Status | | :--- | :--- | :--- | | Reasoning verification (Phase 2 – lawyer HITL) | Autumn 2026 / Spring 2026 | Planned | | All Industrial toolkit Components Development | Spring 2026 | Planned | | Website/App Development | Summer 2026/2027 – Autumn 2027 | Planned | | Research thesis | Summer 2026/2027 – 2027 | Planned |
2. The Detailed Task Breakdown (For Project Planning & Speaker Notes)
This breakdown details the realistic allocation of the remaining phases mapped across the WSU academic calendar.
AUTUMN 2026 (Finalizing Experimental Work & Analysis) Current Semester – Finalizing the data | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phases 1–3 | Finalizing Adaptive Follow-up Experiments (Grok, DeepSeek, ChatGPT) | 3–4 weeks (Concurrent) | | Phase 4 | Researcher analysis of KeyReason (Ratio decidendi) & Further Literature Study | 4 weeks |
MID-YEAR BREAK / EARLY SPRING 2026 (Reasoning Verification) June/July/August – Allowing maximum flexibility for external partners | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 5 | Lawyer verification of KeyReason (Phase 2 – Lawyer HITL) | 4–6 weeks (Extensive buffer for practitioner schedules) | | Phase 6 | Type II Logic Failure quantification | 1 week |
SPRING 2026 (Industrial Translation) August to November – Core development phase | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 7 | All Industrial Toolkit Components Development (8 outputs) | 2+ weeks | | Phase 8 | Practitioner Manual drafting | 2+ weeks |
SUMMER 2026/2027 ONWARDS (Thesis Finalisation) December to mid-2027 – Final synthesis and submission | Phase | Detailed Activity | Realistic Duration | | :--- | :--- | :--- | | Phase 9 | Core thesis writing, comprehensive results synthesis, and industrial toolkit integration | Ongoing (3+ months) | | Phase 10 | Final thesis formatting, defense preparation, review, and official submission | 3–4 weeks |
Total Observations: 960 (fully verified and audited)
Important Note on Statistical Significance
The findings presented in this Executive Summary are descriptive — they reflect observed patterns in the data. While the dataset is comprehensive (960 observations), statistical testing (planned for Phase 1 completion) will determine which observed differences are statistically significant given the sample size of 20 cases per jurisdiction. The descriptive findings provide strong directional evidence and actionable insights, but final conclusions about the robustness of specific comparisons await statistical confirmation.
Headline Findings (Descriptive)
- No single model dominates all jurisdictions – Performance is highly model-dependent and jurisdiction-dependent.
- Jurisdictional bias is real but model-specific – Gemini shows strong US bias; DeepSeek shows extreme UK specialization; ChatGPT and Perplexity are more balanced.
- Clue effect is model-dependent – Most models improve with citation context (+9.1% average), but Grok declines (-3), revealing critical differences in how models use retrieval vs. reasoning.
- "Correct guessing" phenomenon observed – Preliminary analysis suggests models often predict correctly using general principles rather than jurisdiction-specific ratio decidendi (e.g., NG_001: 5/6 correct but used "negligence" instead of "statutory flavour"). This will be systematically verified in Phase 2 reasoning analysis.
- Universal failures reveal systematic gaps – Cases involving procedural law, limitation periods, statutory interpretation, and document hierarchy consistently challenge all models (NG_009: 0/6; NG_020: 0/6; AU_013: 0/6).
- DeepSeek is uniquely UK-specialized – Perfect 20/20 on UK without clues, but near-random on US (9/20) and Australia (9/20).
- Gemini is the most reliable with citation context – 93.8% average accuracy across all four jurisdictions with clues.
- Grok is the most consistent without clues – 83.8% average accuracy across all four jurisdictions without any case identification.
Industrial Translation:
This research further contributes eight industrial governance outputs including a Traceable Accountability three pillar framework, risk classified model selection matrix, domain specific risk map, five phase implementation roadmap, vendor selection scorecard, expanded Duty of Inquiry Checklist, AI Governance Committee charter template, and incident tracking protocol – translating 960 empirical observations into deployable IT management tools for the Nigerian legal sector.
Sections 15–19 · Theoretical contributions, statistical plan, Phase 2, emerging questions and conclusionReport lines 858–1108
SECTION 15: THEORETICAL CONTRIBUTIONS
15.1 Empirical Evidence for "Stochastic Parrot" Hypothesis (Preliminary)
This research offers preliminary empirical support for Dahl et al.'s "stochastic parrot" hypothesis using real legal data. The NG_001 finding – 5/6 models correct but apparently using the wrong reasoning principle – suggests that LLMs may achieve correct outcomes through probabilistic guessing rather than genuine understanding. Full confirmation awaits Phase 2 reasoning analysis.
15.2 Type II Logic Failure as a Metric
You have operationalized a new metric: Type II Logic Failure (correct verdict, wrong reasoning). This distinguishes between models that understand law (Aligned) and models that successfully guess (General-but-Correct) – a crucial distinction for legal practitioners.
15.3 Model-Specific Jurisdictional Bias Taxonomy
You have created the first taxonomy of LLM jurisdictional bias, demonstrating that:
- Gemini = US-specialized
- DeepSeek = UK-specialized
- ChatGPT = Jurisdiction-agnostic
- Perplexity = Commonwealth-preferring (Nigeria/Australia)
- Grok = Most balanced across jurisdictions
- Claude = Moderate US bias with strong UK retrieval
15.4 Translation Framework Methodology
Your Translation Framework for converting substantive liability predictions to procedural appellate outcomes is a methodological innovation that can be adopted by future legal AI benchmarking studies. It ensures fair scoring regardless of output vocabulary.
15.5 Hard Boundary Rule
The Hard Boundary Rule provides a principled method for handling internally inconsistent responses, ensuring scoring consistency and preventing models from benefiting from contradictions.
15.6 Recency Effect Quantification
You have quantified the recency effect in legal LLM performance – a 16.8% accuracy drop for 2025 cases compared to 2015 cases, providing empirical evidence of training data limitations.
15.7 Domain-Specific Performance Mapping
You have created the first detailed mapping of LLM performance across legal domains, identifying systematic weaknesses in procedural law, purposive construction, document hierarchy, and equity.
15.8 Citation Interference Pattern
You have identified a novel phenomenon: citation interference – where providing citation context actively degrades model performance. Grok's consistent decline with clues (-3 overall) is the first documented case of this pattern in legal LLM research.
15.9 Industrial Translation Framework
This research contributes a methodology for translating academic benchmark findings into industrial governance tools – a meta-contribution that enables future researchers to replicate the translation process.
| Academic Finding | Industrial Output | Translation Method | Evidence Section |
|---|---|---|---|
| Model accuracy tables (Sections 3-7) | Risk-classified selection matrix (14.1) | Weighted scoring: Nigerian accuracy 30%, cross-jurisdictional consistency 20%, high-confidence error rate 15% | 7.1-7.2 |
| Type II Logic Failure (12.3) | Duty of Inquiry Checklist (14.4) | Extracted verification items from failure case analysis (NG_001, AU_013) | 12.3 |
| Citation interference pattern (Grok) | Conditional prompting protocol (14.2.5) | Tested across 5 conditions to identify optimal input format | Experiment 1 |
| Recency effect (16.8% drop, 11.2) | "Post-2023" verification rule (14.3) | Operational guardrail with clear implementation threshold | 11.2 |
| Procedural blindness (NG_009) | Mandatory manual verification for limitation questions | Domain-specific risk classification (Critical) | 13.1 |
| Negative Expected Value (DeepSeek) | "NOT RECOMMENDED" procurement designation | EV threshold: < -0.5 triggers automatic restricted status | 3.5, 5.3 |
| RLHF-induced agreeableness (Exp 3) | Adversarial pushback protocol (14.2.5) | Two-step verification: fresh rerun + challenge | Experiment 3 |
| Shadow IT risk (31%) | Phase 1 implementation step (14.6) | Mandatory audit before policy adoption | 14.6 |
This framework ensures that every empirical finding generates an actionable output for legal practitioners and IT managers.
SECTION 16: NEXT STEPS – STATISTICAL TESTING
16.1 Research Questions and Statistical Confirmation (UPDATED)
| Research Question | Initial Finding from Descriptive Data | Primary Statistical Test to Confirm |
|---|---|---|
| RQ1: Which model is most jurisdiction-independent? | Grok leads without clues (83.8%); Gemini leads with clues (93.8%) | Friedman Test + Kendall's W – to determine if differences across jurisdictions are statistically significant |
| RQ1a: Do models show jurisdiction sensitivity? | Yes – DeepSeek: UK 100% vs US 45%; Gemini: US 100% vs Nigeria 60% | Friedman Test by Model – to test if within-model variation across jurisdictions is significant |
| RQ1b: Which model has most consistent accuracy? | ChatGPT: 60-80% range; Grok: 75-90% range | Coefficient of Variation – to quantify consistency relative to mean accuracy |
| RQ2: Does adding clues improve accuracy? | Average improvement +9.1%; Grok declines (-3); Perplexity improves most (+7 in UK) | McNemar's Test – to determine if improvement for each model is statistically significant |
| Narrative A: US vs Nigeria bias | Gemini gap: 40% (US 100% vs Nigeria 60%) | Mann-Whitney U + Cliff's Delta – to test if US-Nigeria gap is significant and measure effect size |
| Narrative B: Commonwealth lineage | DeepSeek shows sharp drop (100% UK → 45% AU); others show mixed trends | Jonckheere-Terpstra Test – to test for ordered trend across UK→AU→NG |
| Calibration: Are models appropriately confident? | DeepSeek: perfect NC calibration but degraded WC; Gemini: excellent both conditions | Reliability Diagrams + ECE – to visualize and quantify calibration quality |
16.2 Why Statistical Testing is Necessary
| Descriptive Finding | What Statistics Will Confirm |
|---|---|
| "Grok scores 18/20 in UK without clues" | Is this significantly better than chance given the 60/40 case distribution? |
| "Gemini shows 40% US-Nigeria gap" | Is this gap statistically significant with n=20 per jurisdiction? |
| "Perplexity improves by +7 in UK" | Is this improvement significant given the sample size? |
| "DeepSeek's UK specialization" | Is the difference between UK and other jurisdictions significant? |
| "Claude regresses with clues in US" | Is this regression statistically meaningful or random variation? |
16.3 Complete Statistical Analysis Workflow
| Step | Analysis | Test | Output | What It Confirms |
|---|---|---|---|---|
| 1 | Descriptive statistics | Mean, Std, CV | Baseline metrics | Initial performance patterns |
| 2 | Overall model comparison | Friedman | p-value for overall difference | Whether models differ significantly overall |
| 3 | If significant | Wilcoxon + Bonferroni | Which models differ | Specific pairwise differences |
| 4 | Clue vs. No Clue | McNemar | p-value per model | Whether clue effect is significant |
| 5 | US vs. Nigeria bias | Mann-Whitney + Cliff's Delta | p-value + effect size | Whether bias gap is significant and its magnitude |
| 6 | Commonwealth trend | Jonckheere-Terpstra | p-value for ordered trend | Whether lineage effect exists |
| 7 | Confidence calibration | Reliability diagrams + ECE | Visual + numeric | Calibration quality |
| 8 | Effect sizes | Cliff's Delta, r, Kendall's W | Magnitude of differences | Practical significance |
| 9 | Correct Guessing Analysis | Researcher + Lawyer | Reasoning Classification | Type II Logic Failure rates |
16.4 Timeline for Statistical Testing
| Phase | Activity | Estimated Duration |
|---|---|---|
| 1 | Data preparation for statistical analysis | 1 day |
| 2 | Descriptive statistics and visualization | 1 day |
| 3 | Friedman tests (overall and by model) | 1 day |
| 4 | Pairwise Wilcoxon with Bonferroni | 1 day |
| 5 | McNemar tests for clue effect | 1 day |
| 6 | Mann-Whitney U for US/Nigeria bias | 1 day |
| 7 | Jonckheere-Terpstra for Commonwealth trend | 1 day |
| 8 | Reliability diagrams and ECE calculation | 1 day |
| 9 | Effect size calculations and synthesis | 1 day |
| 10 | Results interpretation and reporting | 2 days |
Total Estimated Time: 10-12 days
SECTION 17: PHASE 2 – REASONING ANALYSIS
17.1 Two-Step KeyReason Verification
Step 1: Researcher Analysis
- Study each case file to identify ratio decidendi
- Document with page/paragraph references
- Preliminary classification for each LLM using the four-category framework
- Flag potential Type II Logic Failures (General-but-Correct)
Step 2: Lawyer Verification
- Compile complete package for each case:
- Original judgment PDF
- Researcher's extracted ratio decidendi
- Each LLM's stated reasoning
- Preliminary classification
- Legal practitioner provides final verification
- Update master sheet with verified Reasoning Classification
17.2 Reasoning Classification Framework
| Category | Definition | Scoring Impact |
|---|---|---|
| Aligned | LLM reasoning matches jurisdiction-specific principle | Genuine understanding demonstrated |
| General-but-Correct | LLM applies general common law principle that produces same outcome (Type II Logic Failure if jurisdiction-specific principle exists) | Correct outcome but wrong reasoning – signals stochastic parroting |
| General-but-Wrong | LLM applies general principle that would produce wrong outcome (but verdict correct by coincidence) | Correct by luck – unreliable for future cases |
| Misaligned | LLM reasoning contradicts or ignores the actual principle | Fundamental error |
17.3 Cases Selected for Initial Reasoning Analysis
Based on preliminary findings, the following cases are prioritized for Phase 2:
| Case | Reason for Selection |
|---|---|
| NG_001 | Type II Logic Failure candidate (5/6 correct but wrong reasoning – preliminary) |
| AU_013 | Universal failure – understand why all models wrong |
| UK_013 | Mixed performance – understand proprietary estoppel reasoning |
| US_013 | Most difficult recent case – understand recency effect |
| NG_009 | Procedural blindness – understand why models missed limitation point |
| NG_020 | Document hierarchy failure – understand why all models defaulted to wrong intuition |
SECTION 18: NEW RESEARCH QUESTIONS EMERGING FROM THE FINDINGS
The findings presented in this report, while comprehensive and fully addressing the original research questions, open numerous avenues for further investigation. This section catalogs 35 new research questions that emerge directly from the data, organized by thematic area. These questions represent opportunities for future research, potential PhD extensions, and practical investigations that could further enhance understanding of LLM legal reasoning.
18.1 Model Behavior & Performance Questions
- Why does Grok consistently decline with citation clues? Grok is the only model that performs worse with clues in multiple jurisdictions (UK: -2, Australia: -3, Nigeria: 0). Is this because its training data is structured differently such that citation retrieval introduces noise rather than signal? Does Grok over-rely on parametric knowledge and fail to integrate new contextual information?
- What explains DeepSeek's extreme UK specialization? DeepSeek achieves perfect 20/20 on UK without clues but scores only 9/20 on US and Australia. Is its training data disproportionately UK-focused? Would it perform well on other Commonwealth jurisdictions (e.g., Canada, New Zealand, India)?
- Why does Claude regress with clues in the US but improve elsewhere? Claude improves dramatically in UK (+4) and Australia (+5) but regresses in US (-3). Is its US legal knowledge stored differently such that citation retrieval triggers incorrect associations? Does its training data contain systematic errors in US Supreme Court jurisprudence?
- What drives Perplexity's dramatic clue effect? Perplexity improves from 12/20 to 19/20 in UK (+7) but shows no improvement in Australia (0). Is its training data UK-heavy but Australia-light? Does it rely more heavily on retrieval than reasoning, and at what point does this retrieval dependency become a liability?
- Are there "families" of models with similar behavior patterns? Grok and DeepSeek both decline with clues; ChatGPT and Claude both improve consistently. Do models from the same architectural "family" behave similarly? Can we cluster models by performance patterns and infer architectural features from these patterns?
18.2 Domain-Specific Questions
- Why do procedural/limitation cases cause universal failure? NG_009 (statute-bar) achieved 0/6 in No-Clue and only 2/6 in With-Clue. Do LLMs systematically underweight procedural rules in favor of substantive merits? Is this a training data issue (procedural cases underrepresented) or an architectural issue (models biased toward substantive reasoning)?
- Why does document hierarchy (parol evidence) cause universal failure? NG_020 (Atiba v Suberu) achieved 0/6 in No-Clue. Do LLMs lack understanding of integration clauses and document hierarchy? Is this because training data contains more examples of subsequent conduct modifying agreements than formal deed supremacy?
- Why does purposive construction cause universal failure? AU_013 (Ecosse Property Holdings) achieved 0/6 in No-Clue, with all models applying literal textual construction. Do LLMs have a systematic bias toward literal interpretation over purposive construction? Can models be prompted to consider commercial purpose?
- Which legal domains are most and least reliably handled? Domain performance varies dramatically (Contract interpretation 63.9% in Australia, Commercial law 88.9% in Australia). Is there a hierarchy of legal domains by LLM reliability? Which domains are most resistant to improvement with citation clues?
- How do models perform on constitutional vs private law questions? This dataset focuses on private law. Do models perform differently on constitutional questions? Is there a difference between statutory interpretation and common law reasoning?
- Do models perform better on landmark vs obscure cases? Some cases were universally correct, others universally wrong. Is there a correlation between case citation frequency in training data and model accuracy? Do models perform better on cases that appear frequently in legal education materials?
18.3 Temporal Questions
- What is the precise training data cutoff for each model? The sharp performance decline for 2025 cases (61.5% vs 78.3% for 2015 cases) suggests recency effects. Can we estimate each model's training data cutoff date from performance trajectories? Do models have different cutoff dates?
- How does case age interact with jurisdiction? Is the recency effect uniform across jurisdictions or more pronounced in some? Do models perform better on older cases from their "specialized" jurisdictions?
- How will these findings change as models update? Will new model versions show improved performance on previously difficult cases? Does the recency effect persist across model generations? How quickly do models incorporate new jurisprudence?
18.4 Calibration & Confidence Questions
- Why is DeepSeek's calibration perfect in UK but poor elsewhere? DeepSeek's NC Brier Score in UK is 0.0147 (excellent) but degrades in other jurisdictions. Is calibration jurisdiction-specific? Does poor calibration in unfamiliar jurisdictions signal lack of genuine understanding?
- What causes the "overconfidence trap" in some models? DeepSeek expressed 100% confidence on multiple wrong answers (UK_004, UK_015), generating BS of 1.0. Which models are most prone to overconfidence? Does overconfidence correlate with training data gaps?
- Do models know when they don't know? Some models (e.g., Grok) occasionally predicted incorrectly with lower confidence, others (DeepSeek) with high confidence. Which models have the best "metacognitive awareness"? Does calibration improve with citation clues? Can confidence scores be used to triage cases requiring manual review?
18.5 Methodological Questions
- How much did the Translation Framework affect rankings? Without the Translation Framework, three Nigerian scores would have been incorrect. How many scores across all jurisdictions would be affected? Do some models systematically express predictions in substantive terms more than others?
- How often does the Hard Boundary Rule apply? The rule was triggered once in this dataset. Are internally inconsistent responses rare or common? Do certain models produce more inconsistent responses? Should inconsistent responses be treated differently in scoring?
- What is the optimal fact pattern length? Sanitized fact patterns were 75-90 words. Is there an optimal length for legal reasoning? Does information density matter more than absolute length? How much can facts be compressed without losing essential legal content?
18.6 Practitioner-Focused Questions
- Can we create a jurisdiction-specific reliability index? Models have clear jurisdictional strengths and weaknesses. Can we create a "reliability score" for each model in each jurisdiction? How should practitioners weight these scores when choosing models? Should the index be updated as models evolve?
- What is the cost-benefit of using multiple models? Different models excel in different areas. Does using an ensemble of models improve accuracy? What is the optimal ensemble strategy for legal research? How should practitioners reconcile conflicting predictions from different models?
- Can we predict case difficulty for LLMs? Some cases were universally correct, others universally wrong. Can case characteristics (domain, year, procedural posture) predict LLM performance? Is there a "difficulty index" that correlates across models? Can practitioners pre-screen cases for AI reliability based on case metadata?
- How should the Duty of Inquiry Checklist be customized by jurisdiction? Different jurisdictions pose different challenges (e.g., procedural blindness in Nigeria, purposive construction in Australia, document hierarchy in Nigeria). Should verification protocols be jurisdiction-specific? What are the top 3 verification priorities for each jurisdiction?
18.7 Comparative Questions
- How do these models compare to human lawyers? Models achieve 60-100% accuracy depending on condition and jurisdiction. How would experienced lawyers perform on the same task? At what point does AI accuracy exceed typical human performance? Which cases are humans better at than AI?
- How do these models compare to each other statistically? Descriptive differences are clear, but statistical significance needs testing. Which model differences are significant after correction for multiple comparisons? Is the ranking stable across different statistical tests? What sample size is needed to detect meaningful differences?
- Are there "families" of models with similar behavior? Grok and DeepSeek both decline with clues; ChatGPT and Claude both improve. Do models from the same "family" (e.g., both retrieval-augmented) behave similarly? Can we cluster models by performance patterns? Does model architecture predict performance patterns?
18.8 Longitudinal Questions
- How will these findings change as models update? Current snapshot as of early 2026. Will new model versions show improved performance on previously difficult cases? Does the recency effect persist across model generations? How quickly do models incorporate new jurisprudence?
- Can fine-tuning on Nigerian law improve performance? ChatGPT achieves perfect score with clues but Grok leads without clues. How much would fine-tuning on Nigerian case law improve performance? Which model is most amenable to fine-tuning? What is the optimal fine-tuning dataset size and composition?
18.9 Theoretical Questions
- What does "understanding" mean in the context of LLM legal reasoning? Type II Logic Failures (correct verdict, wrong reasoning) show that correctness ≠ understanding. Can we develop a metric for "genuine understanding" that goes beyond outcome prediction? How often do models achieve correct outcomes through incorrect reasoning? What does this tell us about how LLMs represent legal knowledge?
- Is jurisdictional bias a form of epistemic injustice? Some jurisdictions (Nigeria, Australia) are systematically disadvantaged in model performance. Does the performance gap constitute a form of epistemic injustice? Are developing legal systems systematically disadvantaged in AI training data? What are the ethical implications of jurisdictionally biased AI tools?
- What does the clue effect reveal about model architecture? Clue effect varies dramatically by model (Gemini +13, Perplexity +9, Grok -3). Does the clue effect correlate with model architecture (e.g., retrieval-augmented vs pure parametric)? Can we infer architectural differences from performance patterns? How do different models integrate retrieved information with parametric knowledge?
18.10 Practical Implementation Questions
- Can we automate the Duty of Inquiry Checklist? Manual verification is essential but time-consuming. Can we build a tool that automatically verifies citations? Can we flag potential Type II Logic Failures algorithmically? What is the optimal human-AI collaboration model for legal research?
- How should law firms integrate these findings into practice? Clear guidance on model selection and input preparation. What is the best way to communicate findings to practicing lawyers? Should firms develop internal AI usage policies based on these results? How can the Duty of Inquiry Checklist be integrated into existing workflows?
- What is the economic impact of using optimal vs suboptimal models? Model performance varies by 20-40% across jurisdictions. What is the cost of using the wrong model for a given jurisdiction? How many hours of lawyer time could be saved by optimal model selection? What is the ROI of implementing the Duty of Inquiry Checklist?
18.11 Summary: A Research Agenda
These 35 questions represent not merely extensions of this study, but an entire research program. They cluster into several natural research streams:
| Research Stream | Questions | Potential Outputs |
|---|---|---|
| Model Architecture & Behavior | 1-5, 27, 32 | Journal articles on model design |
| Domain-Specific Legal Reasoning | 6-11 | Domain reliability indices |
| Temporal Dynamics | 12-14, 28-29 | Training cutoff estimation methods |
| Calibration & Confidence | 15-17 | Confidence-based reliability tools |
| Methodological Innovation | 18-20 | Improved benchmarking protocols |
| Practical Applications | 21-24, 33-35 | Practitioner tools and guidelines |
| Comparative Studies | 25-26 | Human-AI comparison studies |
| Theoretical Foundations | 30-31 | Epistemic justice in AI |
Priority Recommendations for Immediate Follow-up:
- Investigate Grok's clue-induced decline – This anomalous finding (-3 overall) could reveal important insights about model architecture and retrieval mechanisms.
- Analyze DeepSeek's UK specialization – Understanding this could help predict model performance in other Commonwealth jurisdictions.
- Study procedural blindness – The universal failure on NG_009 has immediate practical implications for legal practitioners.
- Study document hierarchy failure – The universal failure on NG_020 reveals a critical weakness in Nigerian property and banking law contexts.
- Develop jurisdiction-specific reliability indices – Directly useful for practitioners and could become a standard tool.
- Conduct human lawyer comparison study – Essential for contextualizing findings and establishing practical relevance.
SECTION 19: CONCLUSION
19.1 What This Research Has Achieved
| Achievement | Evidence |
|---|---|
| Scale | 960 verified experiments across 4 jurisdictions |
| Rigor | Full audit trail, transparent methodology, documented corrections |
| Originality | First cross-jurisdictional LLM legal study including Nigeria |
| Practical impact | First evidence-based AI guidance for Nigerian legal profession |
| Theoretical contribution | Preliminary empirical support for "stochastic parrot" hypothesis (awaiting Phase 2 confirmation); identification of "citation interference" pattern |
| Methodological innovation | Translation Framework, Hard Boundary Rule, Two-Step Verification |
| Domain mapping | First detailed performance mapping by legal domain |
| Recency quantification | 16.8% accuracy drop for 2025 cases documented |
| Failure pattern identification | Procedural blindness, document hierarchy confusion, purposive construction weaknesses, citation interference |
| Industry Translation: Eight deployable governance outputs | First operational framework for Nigerian legal AI adoption in a regulatory vacuum. |
Industrial Impact Summary:
This research produces nine deployable outputs for the Nigerian legal sector: (1) Traceable Accountability three-pillar framework (Section 14.5); (2) risk-classified model selection matrix with explicit warnings (Section 14.1); (3) Nigerian legal domain risk map (Appendix G); (4) five-phase implementation roadmap (Section 14.6); (5) vendor selection scorecard (Appendix D); (6) expanded Duty of Inquiry Verification Checklist (Section 14.4); (7) AI Governance Committee charter template (Appendix E); (8) incident tracking protocol (Appendix F). These outputs translate 960 empirical observations into actionable IT management tools, enabling Nigerian law firms to deploy LLMs with documented accountability despite the absence of binding AI regulation.
The eight purely empirical outputs are:
| # | Output | Empirical Source |
|---|---|---|
| 1 | Traceable Accountability Three Pillar Framework | Risk patterns from your data |
| 2 | Risk Classified Model Selection Matrix | Accuracy tables (Sections 3–7) + statistical validation |
| 3 | Nigerian Legal Domain Risk Map | Section 10.1 domain accuracy data |
| 4 | Five Phase Implementation Roadmap | Logical sequence from your risk management framework (Section 14.3) – no external statistics |
| 5 | Vendor Selection Scorecard | Weighted criteria from your metrics (Brier, EV, CV, procedural recovery) |
| 6 | Expanded Duty of Inquiry Checklist | Universal failure cases (NG_009, NG_020) + Type II Logic Failure |
| 7 | AI Governance Committee Charter Template | Derived from your risk management framework |
| 8 | Incident Tracking and Remediation Protocol | Derived from error types you observed (hallucinations, procedural blindness, overconfidence) |
19.2 Key Findings Summary (Descriptive) – FINAL CORRECTED
- Grok is most jurisdiction-independent without clues (83.8% across 4 jurisdictions)
- Gemini is most reliable with citation context (93.8% across 4 jurisdictions)
- DeepSeek shows extreme UK specialization (20/20 UK vs 9/20 US) but ⚠️ regresses with clues and exhibits overconfidence on errors
- ChatGPT achieves perfect score on Nigerian law with clues (20/20)
- Perplexity shows largest clue effect (+7 in UK) but remains retrieval-dependent
- Type II Logic Failure (preliminary) – models may be correct for wrong reasons; Phase 2 verification required
- Universal failures reveal systematic gaps in:
- Procedural/limitation law (NG_009)
- Document hierarchy/parol evidence (NG_020)
- Purposive construction (AU_013)
- Citation clues help most models (+9.1% avg) but harm Grok (-3 overall)
- Grok is the only model to consistently decline with clues across multiple jurisdictions (UK -2, Australia -3, Nigeria 0)
- Claude shows largest improvement in Australia (+5) and UK (+4)
- Recency effect – 16.8% accuracy drop for 2025 cases
- Domain-specific weaknesses – procedural law, document hierarchy, purposive construction, equity most challenging
19.3 Next Steps Timeline
| Phase | Activity | Estimated Duration |
|---|---|---|
| 1 | Statistical testing (Friedman, Wilcoxon, McNemar, etc.) | 10-12 days |
| 2 | Researcher analysis of Key Reason (Phase 1) | 2 weeks |
| 3 | Lawyer verification of Key Reason (Phase 2) | 1 week (parallel) |
| 4 | Type II Logic Failure quantification | 3 days |
| 5 | Practitioner Manual drafting | 1 week |
| 6 | Thesis chapter writing (Results & Discussion) | 2 weeks |
| 7 | Industrial toolkit finalisation | Practitioner Manual + 9 outputs | 1 week (parallel to Phase 2) |
19.4 Dual Contribution Statement
This research makes two distinct but interconnected contributions:
Academic Contribution: First systematic cross-jurisdictional LLM legal reasoning benchmark including an African common-law jurisdiction (Nigeria) alongside the US, UK, and Australia. Introduces and operationalises Type II Logic Failure (correct verdict, wrong reasoning) as a metric distinguishing genuine understanding from stochastic parroting. Provides statistical confirmation of model-specific jurisdiction sensitivity, citation interference, and recency effects.
Industrial Contribution: First operational governance framework and practitioner toolkit for responsible AI adoption in a regulatory-vacuum jurisdiction. Delivers nine deployable outputs (enumerated above) that translate empirical findings into IT management tools, procurement criteria, and professional verification standards. Provides evidence-based guidance to the Nigerian Bar Association, NITDA, and Nigerian Law Schools for policy development.
The two contributions are mutually reinforcing: the academic benchmark provides the empirical foundation; the industrial toolkit provides the implementation pathway. Neither is complete without the other.
Future Research Direction – Operationalising Relational AI Governance:
While this study provides an empirically grounded risk management framework, the literature suggests that African ethical values such as Ubuntu (relationality, communal well being) could inform AI governance in culturally resonant ways (Mutswiri et al., 2025; Yilma, 2025). However, Ubuntu’s normative structure remains under specified for direct implementation (Mensah & Van Wynsberghe, 2025). Future research should translate such principles into concrete institutional rules, building on the Traceable Accountability framework developed here. The current study does not claim Ubuntu as a validated contribution.
Report Appendices A–G · Master sheet, sample rows, correction log, governance templates, Nigerian domain risk mapReport lines 1109–1316
APPENDIX A: COMPLETE MASTER SHEET STRUCTURE
| Column | Header | Description |
|---|---|---|
| A | Experiment_ID | Unique identifier: CaseID_Condition_Model |
| B | Case_ID | Original case identifier |
| C | Jurisdiction | Country |
| D | Condition | No Clue or With Clue |
| E | Model | LLM name |
| F | Verdict_Score | 1 = correct, 0 = incorrect |
| G | Confidence | Confidence percentage (0–100) |
| H | Brier_Score | When correct: (Confidence/100 - 1)² |
When incorrect: (Confidence/100 - 0)²
| I | Expected_Value | +Confidence/100 if correct, –Confidence/100 if wrong |
|---|---|---|
| J | Case_Title | Full citation |
| K | Legal_Domain | Area of law |
| L | Year | Judgment year |
| M | Ground_Truth_Reasoning | Researcher-extracted ratio decidendi (with references) |
| N | Reasoning_Classification_Prelim | Researcher's preliminary classification |
| O | Reasoning_Classification_Final | Lawyer's verified classification |
| P | Reasoning_Type | Aligned / General-but-Correct / General-but-Wrong / Misaligned |
| Q | Type_II_Logic_Failure | Yes/No flag for General-but-Correct with jurisdiction-specific principle |
| R | Lawyer_Notes | Additional comments from legal practitioner |
| S | Lawyer_Verified_Date | Date of verification |
APPENDIX B: SAMPLE MASTER SHEET ROWS
| Experiment_ID | Case_ID | Jurisdiction | Condition | Model | Verdict_Score | Confidence | Brier_Score | Expected_Value | Case_Title | Legal_Domain | Year |
|---|---|---|---|---|---|---|---|---|---|---|---|
| NG_001_NoClue_ChatGPT | NG_001 | Nigeria | No Clue | ChatGPT | 1 | 78 | 0.0484 | 0.78 | Damisa v. U.B.A. (2025) | Contract Law | 2025 |
| NG_001_NoClue_Gemini | NG_001 | Nigeria | No Clue | Gemini | 1 | 95 | 0.0025 | 0.95 | Damisa v. U.B.A. (2025) | Contract Law | 2025 |
| NG_001_NoClue_Claude | NG_001 | Nigeria | No Clue | Claude | 1 | 72 | 0.0784 | 0.72 | Damisa v. U.B.A. (2025) | Contract Law | 2025 |
| NG_005_WithClue_DeepSeek | NG_005 | Nigeria | With Clue | DeepSeek | 0 | 100 | 1.0000 | -1.00 | Damisa v. U.B.A. (2025) | Contract Law | 2025 |
| AU_018_WithClue_Grok | AU_018 | Australia | With Clue | Grok | 0 | 75 | 0.5625 | -0.75 | Westpac v. Linside (2020) | Commercial Law | 2020 |
APPENDIX C: CORRECTION LOG
| Date | Section | Original Value | Corrected Value | Reason |
|---|---|---|---|---|
| March 10, 2026 | 6.2 Australia With Clue – Grok total | 13 | 12 | AU_018 Grok score corrected from 1 → 0 (post-hoc verification) |
| March 10, 2026 | 6.3 Australia With Clue Accuracy – Grok | 65.0% | 60.0% | Cascade from above |
| March 10, 2026 | 7.2 Cross-jurisdictional With Clue – Grok total | 65 | 64 | Cascade from above |
| March 10, 2026 | 7.2 Cross-jurisdictional With Clue – Grok avg | 81.3% | 80.0% | Cascade from above |
| March 10, 2026 | 7.3 Clue Effect – Grok improvement | -2 | -3 | Cascade from above |
| March 10, 2026 | 7.3 Mean improvement all models | +9.4% | +9.1% | Cascade from above |
END OF COMPLETE RESEARCH FINDINGS REPORT – FINAL CORRECTED VERSION
All 960 observations verified and audited. Single inconsistency identified and corrected. All totals internally consistent. Statistical testing ready. Phase 2 (Reasoning Analysis) prepared. Practitioner Manual draftable.
APPENDIX D: VENDOR SELECTION SCORECARD
Use this scorecard to evaluate LLM vendors for Nigerian legal practice. Minimum passing score: 75/100.
| Criterion | Weight | Scoring Guide | ChatGPT | Gemini | Claude | Grok | DeepSeek | Perplexity |
|---|---|---|---|---|---|---|---|---|
| Nigerian accuracy with citations | 30% | 100%=30; 95%=28.5; 90%=27; 85%=25.5; 80%=24; <80%=0 | 30 (100%) | 28.5 (95%) | 25.5 (85%) | 24 (80%) | 19.5 (65%) | 24 (80%) |
| Nigerian accuracy without citations | 20% | 80%=16; 75%=15; 70%=14; 65%=13; 60%=12; <60%=0 | 14 (70%) | 12 (60%) | 14 (70%) | 16 (80%) | 10 (50%) | 15 (75%) |
| Cross-jurisdictional consistency | 15% | CV <0.10=15; 0.10-0.15=12; 0.15-0.20=9; >0.20=6; >0.40=0 | 12 (0.1242) | 6 (0.2140) | 9 (0.1695) | 15 (0.0896) | 0 (0.4462) | 12 (0.1091) |
| High-confidence error rate | 15% | <5%=15; 5-10%=12; 10-15%=9; 15-20%=6; >20%=0 | 15 (<5%) | 15 (<5%) | 12 (~8%) | 12 (~8%) | 0 (>20%) | 12 (~8%) |
| Procedural/domain recovery | 10% | Corrected both=10; one=5; none=0 | 10 (both) | 10 (both) | 0 (neither) | 5 (one) | 0 (neither) | 5 (one) |
| Transparency/disclosure | 10% | Full=10; partial=5; none=0 | 5 | 5 | 5 | 5 | 0 | 5 |
| TOTAL SCORE | 100% | 86 | 76.5 | 65.5 | 77 | 29.5 | 73 | |
| Recommendation | APPROVE | APPROVE | CONDITIONAL | APPROVE | REJECT | CONDITIONAL |
Minimum Passing Score: 75/100
Approved Vendors (≥75): ChatGPT (86), Grok (77), Gemini (76.5)
Conditional Approval (65-74): Perplexity (73), Claude (65.5) – requires enhanced verification protocol
Rejected (<65): DeepSeek (29.5) – negative EV, high-confidence error rate, no procedural recovery
Contractual Clauses for Approved Vendors:
1. Vendor must provide quarterly accuracy reports on held-out Nigerian cases
2. Material performance degradation (>15% drop) constitutes breach
3. Vendor must disclose training data updates and recency cutoffs
APPENDIX E: AI GOVERNANCE COMMITTEE CHARTER TEMPLATE
[LAW FIRM NAME] AI GOVERNANCE COMMITTEE CHARTER
Effective Date: _______________
1. Purpose
The AI Governance Committee oversees the responsible adoption, deployment, and monitoring of artificial intelligence tools in legal practice, ensuring compliance with professional obligations and risk management standards derived from empirical benchmarking (Uba, 2026).
2. Membership
- Senior Partner (Chair) – one position
- IT Manager – one position
- Training Officer – one position
- External Ethics Advisor – one position (rotating quarterly)
- Legal Practitioner Representative – one position (rotating annually)
3. Responsibilities
- Evaluate LLM vendors using the Vendor Selection Scorecard (Appendix D)
- Maintain and update the Restricted Model List (models requiring case-by-case approval)
- Review and approve AI Use Policy amendments
- Conduct quarterly tool reviews (accuracy, incident rates, new model versions)
- Review incident reports (see Appendix F) and recommend remediation
- Ensure annual AI literacy training is delivered
- Liaise with NBA and NITDA on policy developments
4. Meeting Schedule
- Quarterly: Full committee review (2 hours)
- Monthly: Subcommittee on incidents (1 hour, as needed)
5. Reporting
- Quarterly report to firm partners (template available)
- Annual governance audit report
- Incident summaries to training officer for CLE materials
6. Decision Authority
- Vendor approval/rejection (requires 2/3 majority)
- Restricted list additions (requires Chair approval)
- Emergency suspension of any AI tool (IT Manager + Chair)
Signatures:
_________________ (Chair) Date: _______________
_________________ (IT Manager) Date: _______________
APPENDIX F: INCIDENT TRACKING TEMPLATE
AI-Related Incident Report Form
| Field | Entry |
|-------|-------|
| Incident ID | AI-[YYYY]-[XXX] |
| Date of incident | _______________ |
| Reporting lawyer | _______________ |
| Model used | ChatGPT / Gemini / Claude / Grok / DeepSeek / Perplexity / Other: ______ |
| Condition at time | With citations / Without citations / Partial clue |
| Case jurisdiction | Nigeria / UK / US / Australia / Other: ______ |
| Legal domain | Contract / Tort / Commercial / Procedural / Property / Equity / Other: ______ |
| Case year | _______________ |
Description of Incident
[What was the AI output? What was the error?]
Error Type (select all that apply)
[ ] Hallucinated citation (non-existent case)
[ ] Wrong jurisdiction applied
[ ] Procedural blindness (missed limitation period)
[ ] Document hierarchy error (parol evidence)
[ ] Purposive construction error (literal interpretation)
[ ] Overconfidence (high confidence on wrong answer)
[ ] Type II Logic Failure (correct verdict, wrong reasoning)
[ ] Other: _______________
Client Impact
[ ] No client impact (caught internally)
[ ] Client advised incorrectly – corrected before filing
[ ] Filed with error – corrected after filing
[ ] Adverse outcome for client
Remediation Actions
[ ] Manual verification performed
[ ] Client notified
[ ] Court notified (if applicable)
[ ] Training updated
[ ] Model added to restricted list (temporary/permanent)
Submitted to AI Governance Committee on: _______________
Committee Resolution: _______________
Follow-up actions: _______________
APPENDIX G: NIGERIAN LEGAL DOMAIN RISK MAP (COLOUR-CODED)
🟢 GREEN – LOW RISK (Safe for AI delegation with routine verification)
Focus: Areas where LLMs demonstrate near-perfect alignment with Nigerian jurisprudence.
| Domain | Accuracy (NC) | Verification Required |
|---|---|---|
| Criminal Law | 100% | Standard outcome check against official reports. |
🟡 YELLOW – MEDIUM RISK (AI use acceptable with standard verification including reasoning)
Focus: General common law principles that transfer well but require "Reasoning Audits."
| Domain | Accuracy (NC) | Verification Required |
|---|---|---|
| Tort Law | 72.0% | Cross-verify outcome + legal reasoning logic. |
| Contract Law (Labour) | 77.8% | Check for Type II Logic Failure (Right answer, wrong reason). |
| Contract Law (Non-land) | 75% (est.) | Verify specific jurisdiction-dependent principles. |
🟠 ORANGE – HIGH RISK (AI use only with mandatory manual verification)
Focus: Highly specialized Nigerian doctrines where models frequently hallucinate or apply US/UK rules.
| Domain | Accuracy (NC) | Verification Required |
|---|---|---|
| Contract Law (Land/Mortgage) | 58.3% | Mandatory manual verification of document hierarchy. |
| Commercial Law (Banking/IP) | 66.7% | Mandatory procedural check for limitation periods. |
| Property/Equity | 58.3% | Manual verification of Nigeria-specific equitable rules. |
🔴 RED – CRITICAL RISK (AI should NOT be used without independent verification)
Focus: Structural failure zones where LLM architecture is "Procedurally Blind."
| Domain | Accuracy (NC) | Verification Required |
|---|---|---|
| Procedural/Limitation Law | 0% (NG_009) | DO NOT RELY ON AI. Manual verification is mandatory. |
| Document Hierarchy/Parol Evid. | 0% (NG_020) | DO NOT RELY ON AI. Manual verification is mandatory. |
One-page Quick Reference for Practitioners
| Colour | Meaning | Action to be Taken |
|---|---|---|
| 🟢 GREEN | Safe to delegate | Use AI; verify the final outcome. |
| 🟡 YELLOW | Use with caution | Verify the outcome AND the underlying reasoning. |
| 🟠 ORANGE | High risk | Mandatory manual verification of specific technical elements. |
| 🔴 RED | Critical risk | DO NOT rely on AI. Perform manual research only. |
This risk map is derived from Section 10.1 (Nigeria) and universal failure cases (Section 13). Update annually based on new model versions.
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Benchmarking report · where everything sits
Map of the full benchmarking report
Every empirical section of the report is placed on the study page it informs. The only text not reproduced is an argumentative essay on the study’s long-term relevance and a research-ethics reflection, which contain no results.
- Executive summary, headline findings and timeline · appendices
- Section 1–2 · Methodology, data integrity and research framework · benchmark method
- Sections 3–6 · Complete score sheets for Nigeria, UK, US and Australia · benchmark accuracy
- Section 7 · Cross-jurisdictional comparison and clue effect by model · citation context
- Sections 8–9 · Model specialisation by jurisdiction and Commonwealth lineage · benchmark accuracy
- Section 10 · Performance by legal domain · benchmark accuracy
- Section 11 · Performance by case year · decision year
- Section 12 · The “correct guessing” observation · preliminary reasoning
- Section 13 · Notable failure cases · confidence
- Section 14 · Recommendations and governance tools for practitioners (planned resources) · next steps
- Sections 15–19 · Theoretical contributions, statistical plan, Phase 2, emerging questions and conclusion · appendices
- Report Appendices A–G · Master sheet, sample rows, correction log, governance templates, Nigerian domain risk map · appendices
- Part 2 · Statistical confirmation of the benchmark (report-only tests) · citation context
- Part 3 · From statistical confirmation to follow-up experiment design · continuation experiments
- Continuation experiments report · Experiments 1–3 in full · continuation experiments
- Phase 2 · Reasoning analysis (researcher key-reason review, awaiting lawyer verification) · preliminary reasoning
- Report Appendix A · Supplementary statistical analyses (logistic regression, ECE, domain chi-square, power) · confidence
- Silent Failure analysis and findings report · confidence
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Mapping between the labels used in this thesis and the original labels
Label Words/source Original label Stage
GLMRS-A 50–100 GLMRS-D (prov.) Third
GLMRS-B 100–150 GLMRS-E (prov.) Third
GLMRS-C 150–300 GLMRS-A Second
GLMRS-D 300–500 GLMRS-B Second
GLMRS-E 500–750 GLMRS-C Second
GLMRS No Clue 150–300 Unchanged First (simultaneous)Thesis Table 12.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Fourth stage, run a, 100–150 words per source: ChatGPT verdicts by summarisation tool
Case GT ChatGPT Claude Copilot Gemini Gemini Notebook Correct /5
NG 001 Dism. D 58% ✓ D 68% ✓ D 70% ✓ D 62% ✓ D 60% ✓ 5
NG 002 Dism. D 51% ✓ D 52% ✓ A 52% × A 51% × A 52% × 2
NG 003 Dism. D 74% ✓ D 76% ✓ D 78% ✓ D 74% ✓ D 72% ✓ 5
NG 004 Dism. D 52% ✓ D 50% ✓ D 51% ✓ D 53% ✓ D 52% ✓ 5
NG 005 Dism. D 57% ✓ D 58% ✓ D 66% ✓ D 57% ✓ D 57% ✓ 5
NG 006 Allow. A 53% ✓ A 50% ✓ A 51% ✓ A 50% ✓ A 51% ✓ 5
NG 007 Dism. D 62% ✓ D 72% ✓ D 76% ✓ D 58% ✓ D 58% ✓ 5
NG 008 Dism. D 67% ✓ D 58% ✓ D 73% ✓ D 69% ✓ D 68% ✓ 5
NG 009 Dism. A 61% × A 56% × A 67% × A 67% × A 64% × 0
NG 010 Dism. D 58% ✓ D 62% ✓ D 61% ✓ D 59% ✓ D 60% ✓ 5
NG 011 Allow. A 82% ✓ A 84% ✓ A 88% ✓ D 88% × A 78% ✓ 4
NG 012 Dism. D 50% ✓ D 50% ✓ D 50% ✓ D 51% ✓ D 51% ✓ 5
NG 013 Allow. A 51% ✓ A 52% ✓ A 52% ✓ A 52% ✓ A 53% ✓ 5
NG 014 Allow. A 54% ✓ A 57% ✓ A 57% ✓ A 54% ✓ A 55% ✓ 5
NG 015 Dism. D 52% ✓ D 53% ✓ D 54% ✓ D 53% ✓ D 52% ✓ 5
NG 016 Dism. D 88% ✓ D 83% ✓ D 86% ✓ D 86% ✓ D 82% ✓ 5
NG 017 Allow. A 52% ✓ A 51% ✓ D 50% × A 50% ✓ A 50% ✓ 4
NG 018 Dism. D 72% ✓ D 66% ✓ A 82% × D 70% ✓ D 67% ✓ 4
NG 019 Allow. A 52% ✓ A 50% ✓ A 58% ✓ D 54% × D 55% × 3
NG 020 Dism. A 53% × A 61% × A 55% × A 63% × A 56% × 0
Correct 18/20 18/20 15/20 15/20 16/20 82/100Thesis Table 12.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Fourth stage, run a, 150–300 words per source: ChatGPT verdicts by summarisation tool
Case GT ChatGPT Claude Copilot Gemini Gemini Notebook Correct /5
NG 001 Dism. D 62% ✓ D 72% ✓ D 72% ✓ D 66% ✓ D 61% ✓ 5
NG 002 Dism. D 52% ✓ D 52% ✓ A 53% × A 52% × A 53% × 2
NG 003 Dism. D 82% ✓ D 81% ✓ D 82% ✓ D 78% ✓ D 78% ✓ 5
NG 004 Dism. D 52% ✓ D 50% ✓ D 51% ✓ D 52% ✓ D 52% ✓ 5
NG 005 Dism. D 61% ✓ D 64% ✓ D 69% ✓ D 59% ✓ D 59% ✓ 5
NG 006 Allow. A 54% ✓ A 63% ✓ A 51% ✓ A 50% ✓ A 54% ✓ 5
NG 007 Dism. D 66% ✓ D 71% ✓ D 79% ✓ D 61% ✓ D 59% ✓ 5
NG 008 Dism. D 70% ✓ D 63% ✓ D 78% ✓ D 72% ✓ D 71% ✓ 5
NG 009 Dism. A 67% × A 59% × A 72% × A 72% × A 67% × 0
NG 010 Dism. D 63% ✓ D 71% ✓ D 64% ✓ D 62% ✓ D 63% ✓ 5
NG 011 Allow. D 91% × D 89% × A 90% ✓ D 92% × A 75% ✓ 2
NG 012 Dism. D 50% ✓ D 50% ✓ D 50% ✓ D 52% ✓ D 51% ✓ 5
NG 013 Allow. A 51% ✓ A 54% ✓ A 54% ✓ A 57% ✓ A 54% ✓ 5
NG 014 Allow. A 55% ✓ A 60% ✓ A 58% ✓ A 56% ✓ A 57% ✓ 5
NG 015 Dism. D 53% ✓ D 54% ✓ D 55% ✓ D 55% ✓ D 53% ✓ 5
NG 016 Dism. D 93% ✓ D 88% ✓ D 89% ✓ D 91% ✓ D 85% ✓ 5
NG 017 Allow. A 52% ✓ A 53% ✓ D 50% × A 51% ✓ A 50% ✓ 4
NG 018 Dism. D 82% ✓ D 72% ✓ A 86% × D 76% ✓ D 72% ✓ 4
NG 019 Allow. A 53% ✓ A 51% ✓ A 59% ✓ D 56% × D 56% × 3
NG 020 Dism. A 55% × A 67% × A 60% × A 68% × A 61% × 0
Correct 17/20 17/20 15/20 15/20 16/20 80/100Thesis Table 12.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Fourth stage, run b (reruns): ChatGPT verdicts with ChatGPT and Claude summaries
100–150 words 150–300 words
Case GT ChatGPT Claude ChatGPT Claude
NG 001 Dism. D 64% ✓ D 62% ✓ D 66% ✓ D 65% ✓
NG 002 Dism. A 52% × D 35% ✓ A 52% × D 35% ✓
NG 003 Dism. D 76% ✓ D 73% ✓ D 82% ✓ D 78% ✓
NG 004 Dism. D 53% ✓ D 36% ✓ D 53% ✓ D 36% ✓
NG 005 Dism. D 67% ✓ D 46% ✓ D 70% ✓ A 57% ×
NG 006 Allow. A 55% ✓ A 28% ✓ A 56% ✓ A 30% ✓
NG 007 Dism. D 79% ✓ D 55% ✓ D 82% ✓ D 64% ✓
NG 008 Dism. D 72% ✓ D 44% ✓ D 77% ✓ D 53% ✓
NG 009 Dism. A 70% × D 43% ✓ A 75% × D 52% ✓
NG 010 Dism. D 61% ✓ D 48% ✓ D 65% ✓ D 54% ✓
NG 011 Allow. A 86% ✓ A 78% ✓ D 92% × D 72% ×
NG 012 Dism. D 51% ✓ D 38% ✓ D 51% ✓ D 38% ✓
NG 013 Allow. A 54% ✓ A 39% ✓ A 56% ✓ A 40% ✓
NG 014 Allow. A 58% ✓ A 41% ✓ A 59% ✓ A 42% ✓
NG 015 Dism. D 55% ✓ D 39% ✓ D 55% ✓ D 39% ✓
NG 016 Dism. D 88% ✓ D 82% ✓ D 92% ✓ D 88% ✓
NG 017 Allow. A 50% ✓ A 31% ✓ A 50% ✓ A 32% ✓
NG 018 Dism. D 66% ✓ D 58% ✓ D 72% ✓ D 67% ✓
NG 019 Allow. A 58% ✓ A 32% ✓ A 61% ✓ A 33% ✓
NG 020 Dism. A 62% × A 40% × A 70% × A 48% ×
Correct 17/20 19/20 16/20 17/20
Confidence values below 50% in the Claude-summary reruns are inconsistent with a stated
binary verdict (see the scoring notes).Thesis Table 12.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Word count of each stored summary (summary text only, exclud- ing the title)
100–150 words 150–300 words
No. Source GPT Claude Copil. Gemini GNB GPT Claude Copil. Gemini GNB
1 Constitution 1999 127 150 117 144 144 213 293 185 255 231
2 Labour Act 136 150 136 145 144 216 293 200 253 200
3 Trade Marks Act 146 149 138 144 144 238 293 209 274 181
4 Land Use Act 141 147 145 141 137 230 296 205 266 185
5 BOFIA 2020 141 146 139 136 n/a∗ 207 299 191 267 n/a∗
6 CBN Cons. Prot. Framework 134 150 136 143 127 216 287 194 277 174
7 Admiralty Jurisdiction Act 141 145 137 138 135 246 295 200 282 155
8 Sheriffs & Civil Process Act 139 150 132 140 148 228 297 192 289 165
9 NAFDAC Act 140 149 128 139 142 219 295 173 281 184
10 Counterfeit & Fake Drugs Act 134 148 134 139 147 226 299 164 286 181
11 Merchant Shipping Act 145 148 131 146 132 222 298 170 288 171
GPT = ChatGPT, Copil. = Microsoft Copilot, GNB = Gemini Notebook. ∗ Gemini Notebook
was supplied with BOFIA but did not produce a usable BOFIA summary at either length.Thesis Table 12.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Benchmark cases and Ground Truth
Case Short title Ground Truth
NG 001 Damisa v. U.B.A. Appeal Dismissed
NG 002 Super Ceramics v. H.E.P. Eng. Appeal Dismissed
NG 003 Moore Associates v. Exphar Appeal Dismissed
NG 004 F.H.A. v. Oyedeji Appeal Dismissed
NG 005 N.Y.S.C. v. Ukachukwu Appeal Dismissed
NG 006 Akaolisa v. Okuma Appeal Allowed
NG 007 Skye Bank v. Adegun Appeal Dismissed
NG 008 Heritage Bank v. Bentworth Appeal Dismissed
NG 009 Ethiopian Airlines v. Polaris Appeal Dismissed
NG 010 Austin Laz v. GTBank Appeal Dismissed
NG 011 Glenyork v. Panalpina Appeal Allowed
NG 012 Olaniran v. Adebayo Appeal Dismissed
NG 013 ACMEL v. First Bank Appeal Allowed
NG 014 Total E&P v. Okwu Appeal Allowed
NG 015 A.B.C. Transport v. Omotoye Appeal Dismissed
NG 016 Barewa Pharm. v. F.R.N. Appeal Dismissed
NG 017 Standard Chartered v. Ameh Appeal Allowed
NG 018 BPS Eng. v. F.R.M.A. Appeal Dismissed
NG 019 Omni Products v. Union Bank Appeal Allowed
NG 020 Atiba Iyalamu v. Suberu Appeal Dismissed
Fourteen cases were dismissed and six were allowed.Thesis Table 12.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Project resources
Resource Specification Cost Status
(AUD)
LLM access Six models, 960 primary interactions and About 300 Expended
the continuation experiments
NWLR Online access Premium Nigerian law reports In kind Secured
Lawyer verification Two practising lawyers for the Phase 4 In kind Secured
reasoning review
Contingency Additional model access and verification 100 Proposed
support if requiredThesis Table 12.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.