Key takeaways
- Lawyer-led reasoning review is in progress.
- Twelve follow-up actions are listed in the thesis.
- Five proposed practitioner resources remain plans.
What was tested & why
Next steps & practitioner resources
Planned resources: Practitioner’s Manual, Prompt Optimisation Playbook, Model Selection Matrix, Duty of Inquiry Checklist and Traceable Governance Framework. Beneficiaries described in the proposal include legal practitioners, courts, policy and governance audiences. No resource is presented here as an endorsed selection guide.
Planned work / awaiting lawyer verification.
Source: Thesis §10.11; proposal Impact Statement · source-reported unless otherwise noted.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 14 · Recommendations and governance tools for practitioners (planned resources)Report lines 703–857
SECTION 14: RECOMMENDATIONS FOR LEGAL PRACTITIONERS
The following recommendations are organised into three industrial tiers:
(1) Immediate actions for individual practitioners (Sections 14.1-14.2);
(2) Organisational governance for law firms (Sections 14.3-14.4); (
3) Procurement and vendor management for IT directors (Sections 14.5-14.6).
14.1 Model Selection Guide: Risk-Stratified Decision Matrix
| Use Case | Recommended Model | Rationale | Risk Tier | EV / Calibration Note |
|---|---|---|---|---|
| Nigerian law (with citations) | ChatGPT | Perfect 20/20 with clue | LOW | EV uniformly positive |
| Nigerian law (blind facts) | Grok | Best unaided (16/20) | MEDIUM | Verify procedural points |
| UK law (without citations) | DeepSeek | Perfect 20/20 without clue | CRITICAL | ⚠️ EV negative on errors; 100% confidence on wrong answers |
| UK law (with citations) | Claude or Perplexity | 19/20 (95%) | LOW | No overconfidence pattern |
| US law (with or without) | Gemini | Perfect 20/20 both | LOW | Avg Brier 0.0025 |
| Australian law (with citations) | Gemini | 18/20 with clue | LOW | Most reliable |
| Procedural/limitation questions | None – verify manually | All failed NG_009 | CRITICAL | Do not rely on AI |
| Document hierarchy/parol ev. | None – verify manually | All failed NG_020 | CRITICAL | Do not rely on AI |
Important Note on Grok: Grok is the only model that consistently declines with citation clues (UK -2, Australia -3, Nigeria 0). For any jurisdiction where citations are available, practitioners should test Grok both with and without citations and compare results. The unaided version may be more reliable.
14.2 Input Preparation Protocol
- Always include full case citation – improves accuracy by average 9.1%
- Include jurisdiction explicitly – essential for models to apply correct legal framework
- Be aware of model-specific behaviors:
- Grok may perform worse with citations – test both conditions and compare
- DeepSeek excels on UK law without citations but regresses with clues – verify both
- Claude improves dramatically with UK citations (+4) but may regress on US (-3)
- Perplexity is retrieval-dependent – requires citations for reliable performance
- For procedural/limitation questions, verify manually – all models failed NG_009
- For document hierarchy/parol evidence questions, verify manually – all models failed NG_020
- For cases after 2023, exercise heightened caution – significant recency effect observed
14.2.5 Model-Specific Prompting Cheat Sheet
| If Using... | With Citations | Without Citations | Never Do | EV Impact |
|---|---|---|---|---|
| ChatGPT | Always include full citation | Acceptable but verify reasoning | Assume US performance = Nigerian performance | Positive |
| Gemini | Gold standard – mandatory | Suboptimal – avoid | Use without jurisdiction specification | Positive |
| Claude | Strong improvement (UK/AU: +4 to +5) | Moderate (60-75%) | Assume US performance transfers | Positive (with caution) |
| Grok | Partial Clue only (year + court level) | Optimal (82.5% combined) | Provide full citation upfront | Negative with full clue |
| DeepSeek | Verify every output (even with clues) | NOT RECOMMENDED for non-UK | Trust high-confidence predictions | Negative (approaching -1.0) |
14.3 Risk Management Framework
| Risk Type | Identified Pattern | Mitigation |
|---|---|---|
| Procedural blindness | NG_009 universal failure | Manually verify limitation/procedural points |
| Document hierarchy confusion | NG_020 universal failure | Manually verify priority of documents |
| Overconfidence | DeepSeek's 100% confident wrong answers | Always verify high-confidence predictions |
| Citation interference | Grok's decline with clues (-3) | Test with and without citations |
| Jurisdiction confusion | DeepSeek's UK/Nigeria gap (20 vs 10) | Never assume cross-jurisdictional transfer |
| Recency gap | US_013 (2025) difficulty | Verify recent cases independently |
| Domain-specific weakness | AU_013 (contract interpretation) | Extra caution with purposive construction cases |
| Type II Logic Failure | NG_001 correct but wrong reasoning (preliminary) | Verify reasoning, not just outcome |
14.4 The Duty of Inquiry Checklist (Draft)
text
DUTY OF INQUIRY CHECKLIST – AI RESEARCH VERIFICATION
Case: _______________ Date: _______________
Case Year: ___________ Jurisdiction: ___________
Before relying on AI-generated research, verify:
VERDICT VERIFICATION
[ ] Does the AI's predicted outcome match independent research?
[ ] Have you checked procedural/limitation points manually?
[ ] For cases after 2023, have you verified with primary sources?
REASONING VERIFICATION
[ ] Does the AI's reasoning match the actual ratio decidendi?
[ ] Has the AI used general principles instead of jurisdiction-specific reasoning?
[ ] Are all citations verifiable in official law reports?
[ ] Does the cited case actually stand for the principle claimed?
[ ] Have you checked for hallucinated citations (cases that do not exist)?
DOMAIN-SPECIFIC CHECKS
[ ] For contract interpretation cases, verify purposive vs literal construction
[ ] For property/equity cases, verify jurisdiction-specific principles
[ ] For procedural questions, verify limitation periods and ouster clauses
[ ] For document hierarchy questions, verify which document governs (deed vs letter)
I confirm that I have verified the above items.
Signature: ____________________
14.5 Traceable Accountability Three-Pillar Framework
This framework translates the empirical findings into a deployable governance architecture for Nigerian law firms, aligned with NIST AI RMF 1.0 (Govern, Map, Measure, Manage).
Pillar 1: Technical Controls
- Vendor selection scorecard (see Appendix D) with weighted criteria: Nigerian accuracy with citations (30%), Nigerian accuracy without citations (20%), cross-jurisdictional consistency (15%), high-confidence error rate (15%), procedural/domain recovery (10%), transparency (10%)
- Mandatory query logging with timestamp, model version, prompt, and output
- Automated citation verification where feasible (integration with NigeriaLII, LawPavilion)
- Restricted model list (e.g., DeepSeek for Nigerian law requires case-by-case partner approval)
Pillar 2: Legal/Professional Controls
- Signed Duty of Inquiry Checklist (Section 14.4) filed with matter documents
- Personal accountability for AI-assisted submissions (signature requirement)
- Mandatory disclosure of AI use to clients (in engagement letters)
- Continuing Legal Education (CLE) requirement: annual AI literacy training
Pillar 3: Organisational Controls
- AI Governance Committee charter (see Appendix E) with membership: senior lawyer (chair), IT manager, training officer, external ethics advisor
- Quarterly tool review and incident tracking (incident template in Appendix F)
- Annual AI literacy training for all legal staff (90-minute module, slide deck available)
- Vendor approval process: any new LLM must pass scorecard minimum (75/100)
Implementation Metrics (from 5-phase roadmap, Section 14.6):
- Month 1: Shadow IT audit complete; 100% staff trained
- Month 2: AI Use Policy adopted; checklists in use
- Month 3: Governance Committee formed
- Month 4: Restricted model list published; approved vendor whitelist established
- Ongoing: Quarterly incident review; annual governance audit
14.6 Five-Phase AI Governance Implementation Roadmap
| Phase | Timeline | Actions | Success Metrics | Responsible Party |
|---|---|---|---|---|
| Phase 1: Awareness | Month 1 | Conduct internal audit of AI tool usage across all staff; mandatory staff training on AI risks and failure modes (NG_009, NG_020 case studies); distribute Duty of Inquiry Checklist draft. | 100% staff trained; inventory of AI tools completed; high-risk models (DeepSeek) identified | IT Manager + Training Officer |
| Phase 2: Policy | Month 2 | Draft and adopt AI Use Policy; finalise mandatory Duty of Inquiry Checklist; issue model selection guidance (based on Section 14.1) | Policy approved by partners; checklists in use; restricted models list published | Managing Partner + IT Manager |
| Phase 3: Governance | Month 3 | Establish AI Governance Committee (charter per Appendix E); first quarterly review; develop vendor selection scorecard | Committee formed; first quarterly report completed; evaluation criteria established | Senior Partner (Chair) |
| Phase 4: Procurement | Month 4 | Evaluate tools against scorecard (minimum passing score 75/100); phase out high-risk tools (DeepSeek); negotiate transparency requirements with vendors | Approved vendor list established; high-risk tools replaced or restricted | IT Manager + Procurement |
| Phase 5: Monitoring | Ongoing | Incident tracking log (template Appendix F); annual governance review; update training and checklists based on new model versions | Incident log maintained; continuous improvement documented; quarterly re-benchmarking on held-out cases | AI Governance Committee |
14.7 Policy and Ecosystem Recommendations
Based on the empirical findings, the following institutional recommendations are made:
For the Nigerian Bar Association (NBA):
- Issue specific guidance on AI use in legal practice (drawing on Section 14.1-14.4)
- Mandatory CLE requirement: 2 hours of AI literacy per biennium
- Adopt the Duty of Inquiry Checklist (Section 14.4) as a professional standard
- Publish warnings about high-risk models (DeepSeek for Nigerian law) with EV justification
For the National Information Technology Development Agency (NITDA):
- Classify legal AI as "High-Risk" under any future AI regulation (following EU AI Act Article 6)
- Mandate transparency requirements: vendors must disclose training data composition by jurisdiction
- Establish testing and certification for legal AI tools using the methodology in this study
- Require incident reporting for AI-related legal errors
For Nigerian Law Schools (Nigerian Law School, university law faculties):
- Integrate AI literacy into curricula including:
- How LLMs work (stochastic parrots, Section 2.1 of literature)
- Documented failure cases (NG_009, NG_020 from Section 13)
- Verification protocols (Section 14.4)
- Domain-specific risk mapping (Appendix G)
For Legal Tech Vendors:
- Disclose training data composition and jurisdictional coverage
- Provide jurisdiction-specific accuracy data (as in Sections 3-6 of this report)
- Implement transparency features: confidence calibration displays, citation verification
- Support Nigerian-specific benchmarks (this study provides the template)
- Provide warnings about known failure modes (procedural blindness, document hierarchy)
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Thesis §10.11 · planned work
Twelve recommended next steps
Complete the reasoning verification (Phase 4). Complete the researcher extraction of the ratio decidendi and the lawyer review of model justifications, and report Type II Logic Failure rates by model, jurisdiction and condition.
Translate the findings for practice (Phase 5). Develop the planned practitioner resources, namely a practitioner’s manual, a prompt optimisation playbook, a model selection matrix, a duty of inquiry checklist and a traceable governance framework, with any recommendations limited to what the evidence supports.
Add a human-lawyer baseline and repeated runs to the crossjurisdiction benchmark. Ask practising lawyers to complete the same prediction task, and rerun the primary conditions so that model–jurisdiction patterns can be separated from run-to-run variation.
Use the highest-observed configuration as a reference set-up. Use Claude’s stored summaries at 100–150 words with ChatGPT as the predictor, and complete a second run for the Microsoft Copilot, Gemini and Gemini Notebook conditions so that all five tools are compared on two runs. Investigate why Gemini Notebook did not produce a usable BOFIA summary, for example by resupplying the source in another file format, and record the outcome.
Related-case experiment. Replace the general legal materials with related, already-decided cases from the same legal domain, using the reference set-up, to test whether case-level context adds to or replaces the general materials.
Generalisation. Test the reference set-up on other legal domains and, if needed, another type of law, and obtain a practising lawyer’s review of the jurisdiction comparison before interpreting cross-country differences.
Replicate the summary-length curve with stored summaries and repeated runs. Store one set of relational summaries per model and length, reuse it across cases, and run each condition several times. Only then can the observed length pattern (65.8%, 65.8%, 67.5%, 65.8% and 64.2%) be attributed to summary length rather than to variation in summary construction or runto-run noise. Further GLMP reruns (runs c and d) would also strengthen the consistency analysis.
Audit the summaries. Lawyer-led, source-grounded review of the stored fourth-stage summaries (faithfulness, omission of exceptions and qualifications, invented content, word-count compliance), starting with sources relevant to the tool-sensitive cases such as BOFIA and the Admiralty Jurisdiction Act, would allow prediction performance to be interpreted alongside representation fidelity [18, 49, 23].
Control chat memory and confidence elicitation. Run every condition in a fresh, unpersonalised session with memory features disabled, record the setting, and define confidence explicitly as the probability (50–100%) that the stated verdict is correct.
Standardise and record the session structure and clue content. Record whether materials are preloaded afresh for each case, and store the exact clue text supplied in each With Clue run.
Audit verdict–reasoning consistency and appellant mapping. Identify errors that arise from mapping a correct legal analysis to the wrong party’s appeal, and test any prompt refinement (for example, stating the appellant explicitly) as a separate, documented condition rather than altering the existing design.
Increase the case and run count. A larger and more balanced set of Allowed and Dismissed cases, with repeated runs, would support confirmatory inference and a calibration analysis (including ECE) that the present sample cannot.