Key takeaways
- Twenty final appellate cases were selected per jurisdiction.
- Sanitised facts were compared with added citation-based context.
- Partial Verdict and Hard Boundary scoring rules were applied.
What was tested & why
How the benchmark was built
The translation framework codes the appellant’s procedural position so an Allowed or Dismissed label refers to the correct appeal. The Partial Verdict Rule scores “allowed/dismissed in part” according to the direction of the partial outcome. The Hard Boundary Rule gives an explicit procedural verdict label precedence over contradictory reasoning. Both scoring rules were confirmed by the researcher after the initial source review.
Phase 4 lawyer verification of reasoning is still in progress; these rules concern verdict scoring, not verified legal reasoning.
Source: Thesis Chapter 3 · source-reported unless otherwise noted.
Method / scoring
From facts to a scored verdict
Qualify the case
Final appellate decisions were filtered through the seven criteria in Table 3.3.
Sanitise the prompt
Party names, dates and identifying details were replaced while the legal issue remained.
Ask both ways
No Clue uses facts alone; With Clue adds citation context. Six models produced a verdict and confidence.
Match the appeal
Translate the model’s answer to the appellant’s position; score correctness and confidence separately.
“You are a distinguished legal scholar and law expert in [Jurisdiction] Law. You are assisting lawyers with legal research to evaluate legal reasoning capabilities.”
Accuracy = correct verdicts / all scored verdicts.
Brier score = mean squared difference between stated probability and binary outcome; lower is better.
ECE = bin-weighted average absolute confidence–accuracy difference; lower is better.
Silent Failure = incorrect verdict with EV < −0.5, equivalent to >50% stated confidence.
Formulas and operational definitions: Thesis §3.10. Translation Framework details were supplied in the website source instructions; both scoring rules were confirmed by the researcher.
Historical report detail · not a thesis result
How substantive answers became appeal labels
The benchmarking report documents a Translation Framework for answers phrased in substantive terms. The scorer first identifies the appellant’s procedural position and then maps the model’s substantive conclusion to Appeal Allowed or Appeal Dismissed. This is not a separate test of whether the model’s reasoning matches the court’s ratio decidendi.
The same report describes a Partial Verdict Rule for outcomes allowed or dismissed in part and a Hard Boundary Rule giving the model’s explicit procedural verdict label priority over conflicting explanation. The researcher confirmed both rules were applied. Reported trigger counts in draft material are not independently verified here.
Historical benchmarking report, scoring-method passage; researcher confirmation. Thesis §§3.5, 3.8, 3.12 controls the reference-outcome and reasoning-review boundaries.
Results / visual evidence
Nigerian extension · selected pooled conditions
Thesis Tables 4.1–4.2, 6.7 and 7.2 · 120 model–case observations per condition; baselines are reused, not recounted.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 1–2 · Methodology, data integrity and research frameworkReport lines 49–201
SECTION 1: METHODOLOGY AND DATA INTEGRITY
1.0 NIST AI RMF Alignment:
This study's methodology and practitioner recommendations are explicitly aligned with the National Institute of Standards and Technology AI Risk Management Framework (AI RMF 1.0). The four NIST functions – Govern, Map, Measure, Manage – correspond respectively to: (1) the governance framework in Section 14.5; (2) the case selection and jurisdictional mapping in Sections 1.5 and 2.2; (3) the scoring metrics (Brier, EV) and statistical validation in Sections 1.1 and Part 2; and (4) the risk management framework and implementation roadmap in Sections 14.3 and 14.6. This alignment gives the practitioner recommendations internationally recognised governance standing.
1.1 Scoring Methodology
| Metric | Formula | Interpretation |
|---|---|---|
| Verdict Score (1/0) | 1 if predicted = ground truth | Binary accuracy |
| Brier Score | (Confidence/100 - Outcome)² | Calibration penalty (lower is better) |
| Expected Value | +Confidence/100 if correct; -Confidence/100 if wrong | Risk-adjusted performance |
Note on Brier Score Calculation: When Verdict_Score = 1 (correct), Brier Score = (Confidence/100 - 1)². When Verdict_Score = 0 (incorrect), Brier Score = (Confidence/100 - 0)². This penalizes overconfidence on errors and underconfidence on correct predictions.
1.2 Translation Framework
Where models expressed predictions in substantive terms ('Liable'/'Not Liable') rather than procedural appellate terms, a structured translation framework was applied based on the identity of the appellant:
| Rule | Appellant | "Liable" maps to | "Not Liable" maps to |
|---|---|---|---|
| Rule 1 | Defendant/Wrongdoer | Appeal Dismissed | Appeal Allowed |
| Rule 2 | Plaintiff/Victim | Appeal Allowed | Appeal Dismissed |
This framework was essential for fairly scoring models (particularly DeepSeek and Perplexity) that frequently expressed predictions in substantive liability terms. The translation was applied case-by-case with reference to the procedural identity of the appellant in each matter.
1.3 Hard Boundary Rule
Where an LLM's response contains an explicit procedural verdict label that contradicts its own substantive prediction, the stated procedural label governs the score. This ensures scoring consistency regardless of internal inconsistencies. This rule was triggered once in the dataset (NG_019 Perplexity With Clue).
1.4 Partial Verdict Rule
For labels in the form 'Appeal Allowed in part' or 'Appeal Dismissed in part', the directional component (Allowed/Dismissed) is compared with the ground truth. If the direction matches, score = 1; otherwise, score = 0. This rule was applied to seven instances in the Australian dataset.
1.5 Input Sanitization Protocol
To ensure scientific rigor and prevent "data contamination" (AI simply remembering the case name), all input data in the Without Clue condition underwent strict sanitization:
- All party names replaced with neutral placeholders ('Party A'/'Party B')
- All business/institutional names replaced with generic descriptions
- All specific dates replaced with placeholders (Month(x) D(i), YYYY)
- No procedural history included – no reference to lower court decisions
- Fact pattern limited to underlying dispute only – does not reveal outcome
- Target length: 75–90 words per fact pattern
SECTION 2: RESEARCH QUESTIONS AND NARRATIVE FRAMEWORK
2.1 Core Research Questions
| Research Question | What It Seeks to Determine | How Descriptive Findings Address It | What Statistical Analysis Will Confirm |
|---|---|---|---|
| RQ1 (Parent): "Which commercially available LLMs demonstrate the most jurisdiction-independent performance across different legal systems?" | Identifies the LLM(s) that maintain the highest and most stable performance when moving between different legal systems (Nigeria, UK, US, Australia) | Without clues: Grok leads (67/80, 83.8%) |
With clues: Gemini leads (75/80, 93.8%)
| Most balanced overall: ChatGPT (range 60-80%) and Grok (range 75-90%) show smallest jurisdictional gaps | Friedman Test + Kendall's W will determine if overall performance differences across jurisdictions are statistically significant, confirming which model truly leads | |
|---|---|---|
| RQ1a: "Do different LLMs show varying degrees of jurisdiction sensitivity?" | Measures whether some models are more affected than others by changes in legal jurisdiction | YES – Extreme variation: |
- DeepSeek: 55% gap (UK 100% → US 45%)
- Gemini: 40% gap (US 100% → Nigeria 60%)
- ChatGPT: Only 20% gap (most jurisdiction-agnostic)
- Grok: 15% gap (most balanced overall)
| Full taxonomy in Section 8.1 | Friedman Test by Model will test whether within-model variation across jurisdictions is statistically significant, quantifying sensitivity magnitude for each model | |
|---|---|---|
| RQ1b: "Which LLM provides the most consistent accuracy across Nigerian, UK, US, and Australian legal systems?" | Identifies the model with the smallest performance variance across all four jurisdictions | Without clues: |
- Grok: 75-90% range (most consistent)
- ChatGPT: 60-80% range
- Perplexity: 60-75% range
With clues:
- Gemini: 90-100% range (most consistent with citations)
- ChatGPT: 70-100% range
- Claude: 75-95% range Coefficient of Variation will quantify consistency relative to mean accuracy, providing a standardized metric for comparing stability across models and conditions
| RQ1c: "Are some LLMs more 'adaptive' (jurisdiction-independent) than others?" | Determines which models can transfer legal reasoning across jurisdictions without performance degradation | Most adaptive (jurisdiction-independent): |
|---|
- ChatGPT (20% range) – consistently 60-80% across all jurisdictions
- Grok (15% range) – strong performance with minimal variation
Least adaptive (jurisdiction-specialized):
- DeepSeek (55% range) – UK specialist, poor elsewhere
- Gemini (40% range) – US specialist
- Claude (30% range) – moderate US bias
| See Section 8.1 for complete taxonomy | Friedman Test + Post-hoc pairwise comparisons will confirm whether adaptivity differences between models are statistically significant | |
|---|---|---|
| RQ2 (Parent): "How should legal practitioners prepare and input documents to maximize LLM prediction accuracy?" | Establishes evidence-based protocols for document preparation and input formatting | Comprehensive guidance in Section 14.2: |
- Always include full case citations (+9.1% avg improvement)
- Specify jurisdiction explicitly
- Model-specific adjustments required (e.g., Grok may worsen with citations)
- 75-90 word sanitized facts sufficient with proper markers Multiple tests across sub-questions (see below)
| RQ2a: "What document preparation methods yield the highest accuracy?" | Identifies specific preparation techniques that improve LLM performance | Primary finding: Including full case citations improves accuracy by average 9.1% (Section 7.3) |
|---|
Secondary: Jurisdiction specification essential – models perform at chance when jurisdiction ambiguous
| Negative finding: Procedural history omission caused universal failure on NG_009 | McNemar's Test will confirm whether the improvement with citations is statistically significant for each model | |
|---|---|---|
| RQ2b: "Should entire documents be input, or summarized/abstracted versions?" | Tests whether document length/completeness affects prediction quality | Controlled experiment design: |
- Without clue condition used 75-90 word sanitized summaries (Section 1.5)
- With clue condition added only citations, not full documents
- Results show summaries + citations sufficient – full text not necessary when key elements present
- Exception: Procedural cases require explicit limitation period statements (NG_009) Logistic regression could isolate effect of input format while controlling for other variables
| RQ2c: "Which specific legal document components (keywords, summaries, full text) optimize reasoning?" | Isolates the most impactful elements of legal documents for LLM reasoning | Component effectiveness ranking (from evidence): |
|---|
1. Case citations/jurisdiction markers – +9.1% improvement
2. Procedural posture statements – essential for limitation cases (NG_009)
3. Fact summaries (75-90 words) – sufficient when combined with citations
4. Document hierarchy clues – required for parol evidence cases (NG_020)
| 5. Keywords alone – insufficient (models need context) | Feature importance analysis could quantify relative contribution of each component | |
|---|---|---|
| RQ2d: "How does document token length affect prediction quality?" | Examines relationship between input length and accuracy | Controlled finding: 75-90 word sanitized summaries produced strong results (60-100% accuracy depending on model/jurisdiction) |
Implication: Length alone less important than content quality – concise summaries with proper markers outperform lengthy documents without jurisdictional context
| Caveat: Some cases (procedural, document hierarchy) require specific elements regardless of length | Correlation analysis between token length and accuracy (across conditions) could quantify effect. For this Project, Noted: token length is not available because the experimental design used fixed length summaries (75 90 words) for the No Clue condition and only added citations for the With Clue condition. This means length variation is minimal and does not need to be analyzed further. | |
|---|---|---|
| RQ3 (Parent): "What are evidence-based best practices for legal practitioners using AI systems across different jurisdictions and case types?" | Synthesizes all findings into actionable guidance for legal professionals | Complete practitioner toolkit in Section 14: |
- 14.1 Model Selection Guide – jurisdiction/condition-specific recommendations
- 14.2 Input Preparation Protocol – 6-point evidence-based workflow
- 14.3 Risk Management Framework – 8 risk patterns with specific mitigations
- 14.4 Duty of Inquiry Checklist – 16 verification items (draft professional standard) Effect size calculations (Cliff's Delta, Cohen's h) will quantify practical significance of recommendations
Reliability diagrams will inform confidence calibration guidance
2.2 Supporting Narratives
| Narrative | Question | How Descriptive Findings Address It | What Statistical Analysis Will Confirm |
|---|---|---|---|
| Narrative A (Data Bias Check): "Do LLMs perform better on jurisdictions they were trained on?" | Tests whether US-centric training creates performance disparities when models are applied to non-US jurisdictions (Nigeria) | Hypothesis confirmed for some models, rejected for others: |
Gemini (strong US bias):
- US: 100%
- Nigeria: 60%
- Gap: 40% – clear evidence of US training bias
DeepSeek (UK bias, not US):
- UK: 100%
- Nigeria: 50%
- Gap: 50% – confirms training data bias, but for UK, not US
ChatGPT (minimal bias):
- US: 65%
- Nigeria: 70%
- Gap: -5% (Nigeria higher!) – challenges hypothesis
| Conclusion: Bias is model-specific, not universal. Some models (Gemini) show expected US bias; others (DeepSeek) show different specialization; some (ChatGPT) show balanced training. | Mann-Whitney U + Cliff's Delta will test if US-Nigeria gap is statistically significant for each model and measure effect size, quantifying practical magnitude of training data bias | |
|---|---|---|
| Narrative B (Legal Lineage Check): "Can LLMs understand shared common law structure?" | Tests whether models can transfer legal reasoning across the Commonwealth family (UK → Australia → Nigeria) or merely memorize jurisdiction-specific patterns | Success scenario (genuine understanding): Would show UK ≈ Australia ≈ Nigeria |
Failure scenario (memorization): Would show UK high, Australia lower, Nigeria lowest
Findings (model-dependent):
DeepSeek (memorization pattern):
- UK: 100%
- Australia: 45%
- Nigeria: 50%
- Conclusion: Memorized UK law, cannot transfer
Perplexity (reverse trend):
- UK: 60%
- Australia: 75%
- Nigeria: 75%
- Conclusion: Different training data composition
Gemini (mixed):
- UK: 75%
- Australia: 75% (stable)
- Nigeria: 60% (drop)
- Conclusion: Understands common law structure but lacks Nigerian-specific knowledge
ChatGPT/Grok (no clear trend):
| • Show no systematic lineage effect – suggests genuine transferable reasoning | Jonckheere-Terpstra Test will test for ordered trend across UK→AU→NG, determining whether a genuine lineage effect exists for any model |
|---|
Interpretation:
- Significant decreasing trend = memorization pattern
- No significant trend = transferable understanding
- Significant increasing trend = reverse lineage (different training focus)
2.3 Industrial Translation Mapping
| Research Question | Industrial Output | Target User |
|---|---|---|
| RQ1 (Most jurisdiction-independent) | Model Selection Matrix (with risk tiers) | Legal practitioners |
| RQ1a (Jurisdiction sensitivity) | Jurisdictional Bias Taxonomy + Bias Mitigation Protocol | IT procurement |
| RQ1b (Consistency) | Consistency Rankings for vendor evaluation | Risk managers |
| RQ2a (Document preparation) | Input Preparation Protocol | Legal researchers |
| RQ2a (Citation effect) | Model-Specific Prompting Cheat Sheet | Legal practitioners |
| RQ3 (Best practices) | Traceable Accountability Three-Pillar Framework | Law firm partners, IT directors |
| Narrative A (US vs Nigeria bias) | Vendor Disclosure Requirement (contractual clause) | IT procurement |
| Narrative B (Commonwealth lineage) | Transferable reasoning confidence guide | Legal practitioners |
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Case qualification funnel
~630–1,050 decisions
~200–320 decisions
80 cases · 20 per jurisdiction
Thesis Figure 3.1. Approximate screening ranges, not exact case totals.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Phases of the CrossLaw project and their status
Phase Content Status
Phase 1: Dataset Jurisdiction selection, case sourcing, Completed
construction qualification, outcome extraction and
sanitisation
Phase 2: Experimen- Six models, 80 cases, No Clue and Completed
tation With Clue conditions (960 interactions)
Phase 3: Analysis Descriptive, calibration and inferential Completed
analysis, and three continuation experiments
Phase 4: Reasoning Researcher extraction of the ratio decidendi In progress
verification and lawyer verification of model reasoning
Phase 5: Industrial Practitioner resources derived from the Planned
translation findings
Phase 6: Thesis and Final reporting and dissemination Planned
publication
The Nigerian extension (Part II, Chapters 5 to 9) builds on the Phase 2 dataset and was
conducted after Phases 1 to 3.Thesis Table 3.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Primary case sources by jurisdiction
Jurisdiction Primary source(s)
Nigeria NWLR Online, with premium access through two Nigerian
law-firm libraries
Australia AustLII, High Court of Australia decisions
United States Oyez, cross-referenced with the Supreme Court Database
United Kingdom UK Supreme Court archive and BAILIIThesis Table 3.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Seven-criterion CrossLaw case qualification framework
Criterion Operational require- Purpose in the benchmark
ment
Recency Decision between 2015 and To focus on contemporary appellate
2026 law.
Legal area Civil, commercial, regula- To avoid combining materially
tory and related appellate different task types, particularly
disputes criminal or constitutional cases.
Appellate level Supreme Court or final ap- To use authoritative final outcomes.
pellate court
Binary outcome A definitive Appeal Al- To make predictions objectively
lowed or Appeal Dismissed scorable as correct or incorrect.
result
Logical diversity Cases with non-similar rea- To avoid a benchmark dominated by
soning paths repeated applications of one rule.
Subject-matter At least five legal domains To test performance across different
breadth per jurisdiction areas of law.
Proof standard An identifiable condition To retain a meaningful legal decision
precedent or evidentiary problem rather than a superficial
threshold labelling task.Thesis Table 3.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Case sanitisation protocol
Component Treatment
Party names Replaced with neutral placeholders such as Party A and Party
B
Institutional names Converted to generic descriptions where appropriate
Dates Removed or replaced with placeholders
Procedural history Removed where not necessary to the prediction task
Fact-pattern length Standardised to approximately 75–90 words
Legal issue Retained
Judicial outcome Withheld from the modelsThesis Table 3.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Prompt architecture of the primary CrossLaw experiment
Component What was supplied Why it was included
System instruc- A fixed legal-research role with To state which legal system
tion the relevant jurisdiction inserted applied while keeping the role
consistent.
Case-specific in- Sanitised facts under No Clue, To create the controlled
put and the same facts plus citation difference between the two
context under With Clue conditions.
Prediction task Predict the most likely To produce a binary verdict
appellate outcome under the that can be scored.
applicable law
Required re- Verdict, confidence from 0% to To collect comparable verdict
sponse 100% and a brief legal and confidence fields and retain
justification (the fuller template reasoning for later legal review.
also requested precedent
citations)Thesis Table 3.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Inferential and supplementary methods used in the primary benchmark
Method Purpose
McNemar’s test Paired binary comparison of the same model on the same
cases under No Clue and With Clue [44]
Cochran’s Q test Comparison of more than two related binary proportions
on matched cases [11]
Chi-square test Association between correctness and a category such as
jurisdiction or legal domain
Jonckheere–Terpstra test Test of the prespecified ordering high-resource,
mid-resource, low-resource [36]
Mann–Whitney U test Two-group distributional comparisons with a measure of
with Cliff’s delta direction and magnitude
Friedman test with Comparison of multiple models on the same cases
Kendall’s W
Logistic regression Contribution of model and prompting condition to
correctness
Power analysis Adequacy of the sample for detecting medium-sized effects
[13]
The principal significant results are reported in Table 4.12. The Friedman, logistic regression
and power analyses formed part of the analysis plan, but their numerical results are not part of
the consolidated records reported here.Thesis Table 3.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Design of the continuation experiments
Experiment Question Approach
EQ1: Grok cita- Why did Grok’s accu- 40 Australian and United Kingdom
tion sensitivity racy decline when cita- cases retested under five formats:
tion context was added? No Clue, With Clue, Late Reveal,
Partial Clue and Jurisdiction Only.
EQ2: Did the 20/20 United Examination of whether responses
DeepSeek’s Kingdom No Clue result reproduced the original case citation
United Kingdom reflect transferable per- from sanitised facts, comparison with
result formance or possible case other jurisdictions, and a rerun of the 20
familiarity? United Kingdom cases.
EQ3: ChatGPT Could alternative Selected errors retested under
prompt sensitiv- prompts change se- Adversarial Pushback, Jurisdiction
ity lected Australian and Masking and Confidence Penalty
United States errors? prompts.Thesis Table 3.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.