CrossLaw/Part I · Benchmark/How the benchmark was built
04 / 26 · Part I · BenchmarkThesis Chapter 3

80 cases, four jurisdictions, six models, two conditions.

How the benchmark was built

Key takeaways

  • Twenty final appellate cases were selected per jurisdiction.
  • Sanitised facts were compared with added citation-based context.
  • Partial Verdict and Hard Boundary scoring rules were applied.

What was tested & why

How the benchmark was built

The translation framework codes the appellant’s procedural position so an Allowed or Dismissed label refers to the correct appeal. The Partial Verdict Rule scores “allowed/dismissed in part” according to the direction of the partial outcome. The Hard Boundary Rule gives an explicit procedural verdict label precedence over contradictory reasoning. Both scoring rules were confirmed by the researcher after the initial source review.

Evidence boundary

Phase 4 lawyer verification of reasoning is still in progress; these rules concern verdict scoring, not verified legal reasoning.

Source: Thesis Chapter 3 · source-reported unless otherwise noted.

Method / scoring

From facts to a scored verdict

01 / Source

Qualify the case

Final appellate decisions were filtered through the seven criteria in Table 3.3.

02 / Input

Sanitise the prompt

Party names, dates and identifying details were replaced while the legal issue remained.

03 / Test

Ask both ways

No Clue uses facts alone; With Clue adds citation context. Six models produced a verdict and confidence.

04 / Score

Match the appeal

Translate the model’s answer to the appellant’s position; score correctness and confidence separately.

“You are a distinguished legal scholar and law expert in [Jurisdiction] Law. You are assisting lawyers with legal research to evaluate legal reasoning capabilities.”
System instruction · Thesis §3.8

Accuracy = correct verdicts / all scored verdicts.

Brier score = mean squared difference between stated probability and binary outcome; lower is better.

ECE = bin-weighted average absolute confidence–accuracy difference; lower is better.

Silent Failure = incorrect verdict with EV < −0.5, equivalent to >50% stated confidence.

Formulas and operational definitions: Thesis §3.10. Translation Framework details were supplied in the website source instructions; both scoring rules were confirmed by the researcher.

Historical report detail · not a thesis result

How substantive answers became appeal labels

The benchmarking report documents a Translation Framework for answers phrased in substantive terms. The scorer first identifies the appellant’s procedural position and then maps the model’s substantive conclusion to Appeal Allowed or Appeal Dismissed. This is not a separate test of whether the model’s reasoning matches the court’s ratio decidendi.

The same report describes a Partial Verdict Rule for outcomes allowed or dismissed in part and a Hard Boundary Rule giving the model’s explicit procedural verdict label priority over conflicting explanation. The researcher confirmed both rules were applied. Reported trigger counts in draft material are not independently verified here.

Historical benchmarking report, scoring-method passage; researcher confirmation. Thesis §§3.5, 3.8, 3.12 controls the reference-outcome and reasoning-review boundaries.

Results / visual evidence

Data visualization · source-reported

Nigerian extension · selected pooled conditions

Accuracy— 70% always-Dismissed reference
Original No Clue
67.5%
Simultaneous full materials
72.5%
Preloaded full · run a
75.0%
Preloaded full · run b
74.2%
Original With Clue
84.2%

Thesis Tables 4.1–4.2, 6.7 and 7.2 · 120 model–case observations per condition; baselines are reused, not recounted.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Section 1–2 · Methodology, data integrity and research frameworkReport lines 49–201

SECTION 1: METHODOLOGY AND DATA INTEGRITY

1.0 NIST AI RMF Alignment:

This study's methodology and practitioner recommendations are explicitly aligned with the National Institute of Standards and Technology AI Risk Management Framework (AI RMF 1.0). The four NIST functions – Govern, Map, Measure, Manage – correspond respectively to: (1) the governance framework in Section 14.5; (2) the case selection and jurisdictional mapping in Sections 1.5 and 2.2; (3) the scoring metrics (Brier, EV) and statistical validation in Sections 1.1 and Part 2; and (4) the risk management framework and implementation roadmap in Sections 14.3 and 14.6. This alignment gives the practitioner recommendations internationally recognised governance standing.

1.1 Scoring Methodology

MetricFormulaInterpretation
Verdict Score (1/0)1 if predicted = ground truthBinary accuracy
Brier Score(Confidence/100 - Outcome)²Calibration penalty (lower is better)
Expected Value+Confidence/100 if correct; -Confidence/100 if wrongRisk-adjusted performance

Note on Brier Score Calculation: When Verdict_Score = 1 (correct), Brier Score = (Confidence/100 - 1)². When Verdict_Score = 0 (incorrect), Brier Score = (Confidence/100 - 0)². This penalizes overconfidence on errors and underconfidence on correct predictions.

1.2 Translation Framework

Where models expressed predictions in substantive terms ('Liable'/'Not Liable') rather than procedural appellate terms, a structured translation framework was applied based on the identity of the appellant:

RuleAppellant"Liable" maps to"Not Liable" maps to
Rule 1Defendant/WrongdoerAppeal DismissedAppeal Allowed
Rule 2Plaintiff/VictimAppeal AllowedAppeal Dismissed

This framework was essential for fairly scoring models (particularly DeepSeek and Perplexity) that frequently expressed predictions in substantive liability terms. The translation was applied case-by-case with reference to the procedural identity of the appellant in each matter.

1.3 Hard Boundary Rule

Where an LLM's response contains an explicit procedural verdict label that contradicts its own substantive prediction, the stated procedural label governs the score. This ensures scoring consistency regardless of internal inconsistencies. This rule was triggered once in the dataset (NG_019 Perplexity With Clue).

1.4 Partial Verdict Rule

For labels in the form 'Appeal Allowed in part' or 'Appeal Dismissed in part', the directional component (Allowed/Dismissed) is compared with the ground truth. If the direction matches, score = 1; otherwise, score = 0. This rule was applied to seven instances in the Australian dataset.

1.5 Input Sanitization Protocol

To ensure scientific rigor and prevent "data contamination" (AI simply remembering the case name), all input data in the Without Clue condition underwent strict sanitization:

  • All party names replaced with neutral placeholders ('Party A'/'Party B')
  • All business/institutional names replaced with generic descriptions
  • All specific dates replaced with placeholders (Month(x) D(i), YYYY)
  • No procedural history included – no reference to lower court decisions
  • Fact pattern limited to underlying dispute only – does not reveal outcome
  • Target length: 75–90 words per fact pattern

SECTION 2: RESEARCH QUESTIONS AND NARRATIVE FRAMEWORK

2.1 Core Research Questions

Research QuestionWhat It Seeks to DetermineHow Descriptive Findings Address ItWhat Statistical Analysis Will Confirm
RQ1 (Parent): "Which commercially available LLMs demonstrate the most jurisdiction-independent performance across different legal systems?"Identifies the LLM(s) that maintain the highest and most stable performance when moving between different legal systems (Nigeria, UK, US, Australia)Without clues: Grok leads (67/80, 83.8%)

With clues: Gemini leads (75/80, 93.8%)

Most balanced overall: ChatGPT (range 60-80%) and Grok (range 75-90%) show smallest jurisdictional gapsFriedman Test + Kendall's W will determine if overall performance differences across jurisdictions are statistically significant, confirming which model truly leads
RQ1a: "Do different LLMs show varying degrees of jurisdiction sensitivity?"Measures whether some models are more affected than others by changes in legal jurisdictionYES – Extreme variation:
  • DeepSeek: 55% gap (UK 100% → US 45%)
  • Gemini: 40% gap (US 100% → Nigeria 60%)
  • ChatGPT: Only 20% gap (most jurisdiction-agnostic)
  • Grok: 15% gap (most balanced overall)
Full taxonomy in Section 8.1Friedman Test by Model will test whether within-model variation across jurisdictions is statistically significant, quantifying sensitivity magnitude for each model
RQ1b: "Which LLM provides the most consistent accuracy across Nigerian, UK, US, and Australian legal systems?"Identifies the model with the smallest performance variance across all four jurisdictionsWithout clues:
  • Grok: 75-90% range (most consistent)
  • ChatGPT: 60-80% range
  • Perplexity: 60-75% range

With clues:

  • Gemini: 90-100% range (most consistent with citations)
  • ChatGPT: 70-100% range
  • Claude: 75-95% range Coefficient of Variation will quantify consistency relative to mean accuracy, providing a standardized metric for comparing stability across models and conditions
RQ1c: "Are some LLMs more 'adaptive' (jurisdiction-independent) than others?"Determines which models can transfer legal reasoning across jurisdictions without performance degradationMost adaptive (jurisdiction-independent):
  • ChatGPT (20% range) – consistently 60-80% across all jurisdictions
  • Grok (15% range) – strong performance with minimal variation

Least adaptive (jurisdiction-specialized):

  • DeepSeek (55% range) – UK specialist, poor elsewhere
  • Gemini (40% range) – US specialist
  • Claude (30% range) – moderate US bias
See Section 8.1 for complete taxonomyFriedman Test + Post-hoc pairwise comparisons will confirm whether adaptivity differences between models are statistically significant
RQ2 (Parent): "How should legal practitioners prepare and input documents to maximize LLM prediction accuracy?"Establishes evidence-based protocols for document preparation and input formattingComprehensive guidance in Section 14.2:
  • Always include full case citations (+9.1% avg improvement)
  • Specify jurisdiction explicitly
  • Model-specific adjustments required (e.g., Grok may worsen with citations)
  • 75-90 word sanitized facts sufficient with proper markers Multiple tests across sub-questions (see below)
RQ2a: "What document preparation methods yield the highest accuracy?"Identifies specific preparation techniques that improve LLM performancePrimary finding: Including full case citations improves accuracy by average 9.1% (Section 7.3)

Secondary: Jurisdiction specification essential – models perform at chance when jurisdiction ambiguous

Negative finding: Procedural history omission caused universal failure on NG_009McNemar's Test will confirm whether the improvement with citations is statistically significant for each model
RQ2b: "Should entire documents be input, or summarized/abstracted versions?"Tests whether document length/completeness affects prediction qualityControlled experiment design:
  • Without clue condition used 75-90 word sanitized summaries (Section 1.5)
  • With clue condition added only citations, not full documents
  • Results show summaries + citations sufficient – full text not necessary when key elements present
  • Exception: Procedural cases require explicit limitation period statements (NG_009) Logistic regression could isolate effect of input format while controlling for other variables
RQ2c: "Which specific legal document components (keywords, summaries, full text) optimize reasoning?"Isolates the most impactful elements of legal documents for LLM reasoningComponent effectiveness ranking (from evidence):

1. Case citations/jurisdiction markers – +9.1% improvement

2. Procedural posture statements – essential for limitation cases (NG_009)

3. Fact summaries (75-90 words) – sufficient when combined with citations

4. Document hierarchy clues – required for parol evidence cases (NG_020)

5. Keywords alone – insufficient (models need context)Feature importance analysis could quantify relative contribution of each component
RQ2d: "How does document token length affect prediction quality?"Examines relationship between input length and accuracyControlled finding: 75-90 word sanitized summaries produced strong results (60-100% accuracy depending on model/jurisdiction)

Implication: Length alone less important than content quality – concise summaries with proper markers outperform lengthy documents without jurisdictional context

Caveat: Some cases (procedural, document hierarchy) require specific elements regardless of lengthCorrelation analysis between token length and accuracy (across conditions) could quantify effect. For this Project, Noted: token length is not available because the experimental design used fixed length summaries (75 90 words) for the No Clue condition and only added citations for the With Clue condition. This means length variation is minimal and does not need to be analyzed further.
RQ3 (Parent): "What are evidence-based best practices for legal practitioners using AI systems across different jurisdictions and case types?"Synthesizes all findings into actionable guidance for legal professionalsComplete practitioner toolkit in Section 14:
  • 14.1 Model Selection Guide – jurisdiction/condition-specific recommendations
  • 14.2 Input Preparation Protocol – 6-point evidence-based workflow
  • 14.3 Risk Management Framework – 8 risk patterns with specific mitigations
  • 14.4 Duty of Inquiry Checklist – 16 verification items (draft professional standard) Effect size calculations (Cliff's Delta, Cohen's h) will quantify practical significance of recommendations

Reliability diagrams will inform confidence calibration guidance

2.2 Supporting Narratives

NarrativeQuestionHow Descriptive Findings Address ItWhat Statistical Analysis Will Confirm
Narrative A (Data Bias Check): "Do LLMs perform better on jurisdictions they were trained on?"Tests whether US-centric training creates performance disparities when models are applied to non-US jurisdictions (Nigeria)Hypothesis confirmed for some models, rejected for others:

Gemini (strong US bias):

  • US: 100%
  • Nigeria: 60%
  • Gap: 40% – clear evidence of US training bias

DeepSeek (UK bias, not US):

  • UK: 100%
  • Nigeria: 50%
  • Gap: 50% – confirms training data bias, but for UK, not US

ChatGPT (minimal bias):

  • US: 65%
  • Nigeria: 70%
  • Gap: -5% (Nigeria higher!) – challenges hypothesis
Conclusion: Bias is model-specific, not universal. Some models (Gemini) show expected US bias; others (DeepSeek) show different specialization; some (ChatGPT) show balanced training.Mann-Whitney U + Cliff's Delta will test if US-Nigeria gap is statistically significant for each model and measure effect size, quantifying practical magnitude of training data bias
Narrative B (Legal Lineage Check): "Can LLMs understand shared common law structure?"Tests whether models can transfer legal reasoning across the Commonwealth family (UK → Australia → Nigeria) or merely memorize jurisdiction-specific patternsSuccess scenario (genuine understanding): Would show UK ≈ Australia ≈ Nigeria

Failure scenario (memorization): Would show UK high, Australia lower, Nigeria lowest

Findings (model-dependent):

DeepSeek (memorization pattern):

  • UK: 100%
  • Australia: 45%
  • Nigeria: 50%
  • Conclusion: Memorized UK law, cannot transfer

Perplexity (reverse trend):

  • UK: 60%
  • Australia: 75%
  • Nigeria: 75%
  • Conclusion: Different training data composition

Gemini (mixed):

  • UK: 75%
  • Australia: 75% (stable)
  • Nigeria: 60% (drop)
  • Conclusion: Understands common law structure but lacks Nigerian-specific knowledge

ChatGPT/Grok (no clear trend):

• Show no systematic lineage effect – suggests genuine transferable reasoningJonckheere-Terpstra Test will test for ordered trend across UK→AU→NG, determining whether a genuine lineage effect exists for any model

Interpretation:

  • Significant decreasing trend = memorization pattern
  • No significant trend = transferable understanding
  • Significant increasing trend = reverse lineage (different training focus)

2.3 Industrial Translation Mapping

Research QuestionIndustrial OutputTarget User
RQ1 (Most jurisdiction-independent)Model Selection Matrix (with risk tiers)Legal practitioners
RQ1a (Jurisdiction sensitivity)Jurisdictional Bias Taxonomy + Bias Mitigation ProtocolIT procurement
RQ1b (Consistency)Consistency Rankings for vendor evaluationRisk managers
RQ2a (Document preparation)Input Preparation ProtocolLegal researchers
RQ2a (Citation effect)Model-Specific Prompting Cheat SheetLegal practitioners
RQ3 (Best practices)Traceable Accountability Three-Pillar FrameworkLaw firm partners, IT directors
Narrative A (US vs Nigeria bias)Vendor Disclosure Requirement (contractual clause)IT procurement
Narrative B (Commonwealth lineage)Transferable reasoning confidence guideLegal practitioners

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.

Figure 3.1 · native recreation

Case qualification funnel

Screened
~630–1,050 decisions
Seven-criterion qualification
~200–320 decisions
Purposive selection
80 cases · 20 per jurisdiction

Thesis Figure 3.1. Approximate screening ranges, not exact case totals.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 3.1 · source-reported

Phases of the CrossLaw project and their status

Phase                       Content                                         Status

Phase 1:         Dataset    Jurisdiction selection, case sourcing,          Completed
construction                qualification, outcome extraction and
                            sanitisation
Phase 2: Experimen-         Six models, 80 cases, No Clue and               Completed
tation                      With Clue conditions (960 interactions)
Phase 3: Analysis           Descriptive, calibration and inferential        Completed
                            analysis, and three continuation experiments
Phase 4: Reasoning          Researcher extraction of the ratio decidendi    In progress
verification                and lawyer verification of model reasoning
Phase 5:       Industrial   Practitioner resources derived from the         Planned
translation                 findings
Phase 6: Thesis and         Final reporting and dissemination               Planned
publication
   The Nigerian extension (Part II, Chapters 5 to 9) builds on the Phase 2 dataset and was
                               conducted after Phases 1 to 3.

Thesis Table 3.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.2 · source-reported

Primary case sources by jurisdiction

Jurisdiction          Primary source(s)

Nigeria               NWLR Online, with premium access through two Nigerian
                      law-firm libraries
Australia             AustLII, High Court of Australia decisions
United States         Oyez, cross-referenced with the Supreme Court Database
United Kingdom        UK Supreme Court archive and BAILII

Thesis Table 3.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.3 · source-reported

Seven-criterion CrossLaw case qualification framework

Criterion            Operational          require-    Purpose in the benchmark
                     ment

Recency              Decision between 2015 and        To focus on contemporary appellate
                     2026                             law.
Legal area           Civil, commercial, regula-       To avoid combining materially
                     tory and related appellate       different task types, particularly
                     disputes                         criminal or constitutional cases.
Appellate level      Supreme Court or final ap-       To use authoritative final outcomes.
                     pellate court
Binary outcome       A definitive Appeal Al-          To make predictions objectively
                     lowed or Appeal Dismissed        scorable as correct or incorrect.
                     result
Logical diversity    Cases with non-similar rea-      To avoid a benchmark dominated by
                     soning paths                     repeated applications of one rule.
Subject-matter       At least five legal domains      To test performance across different
breadth              per jurisdiction                 areas of law.
Proof standard       An identifiable condition        To retain a meaningful legal decision
                     precedent   or     evidentiary   problem rather than a superficial
                     threshold                        labelling task.

Thesis Table 3.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.4 · source-reported

Case sanitisation protocol

Component              Treatment

Party names            Replaced with neutral placeholders such as Party A and Party
                       B
Institutional names    Converted to generic descriptions where appropriate
Dates                  Removed or replaced with placeholders
Procedural history     Removed where not necessary to the prediction task
Fact-pattern length    Standardised to approximately 75–90 words
Legal issue            Retained
Judicial outcome       Withheld from the models

Thesis Table 3.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.5 · source-reported

Prompt architecture of the primary CrossLaw experiment

Component             What was supplied                    Why it was included

System     instruc-   A fixed legal-research role with     To state which legal system
tion                  the relevant jurisdiction inserted   applied while keeping the role
                                                           consistent.
Case-specific in-     Sanitised facts under No Clue,       To create the controlled
put                   and the same facts plus citation     difference between the two
                      context under With Clue              conditions.
Prediction task       Predict the most likely              To produce a binary verdict
                      appellate outcome under the          that can be scored.
                      applicable law
Required        re-   Verdict, confidence from 0% to       To collect comparable verdict
sponse                100% and a brief legal               and confidence fields and retain
                      justification (the fuller template   reasoning for later legal review.
                      also requested precedent
                      citations)

Thesis Table 3.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.6 · source-reported

Inferential and supplementary methods used in the primary benchmark

Method                           Purpose

McNemar’s test                   Paired binary comparison of the same model on the same
                                 cases under No Clue and With Clue [44]
Cochran’s Q test                 Comparison of more than two related binary proportions
                                 on matched cases [11]
Chi-square test                  Association between correctness and a category such as
                                 jurisdiction or legal domain
Jonckheere–Terpstra test         Test of the prespecified ordering high-resource,
                                 mid-resource, low-resource [36]
Mann–Whitney          U   test   Two-group distributional comparisons with a measure of
with Cliff’s delta               direction and magnitude
Friedman       test       with   Comparison of multiple models on the same cases
Kendall’s W
Logistic regression              Contribution of model and prompting condition to
                                 correctness
Power analysis                   Adequacy of the sample for detecting medium-sized effects
                                 [13]
 The principal significant results are reported in Table 4.12. The Friedman, logistic regression
and power analyses formed part of the analysis plan, but their numerical results are not part of
                             the consolidated records reported here.

Thesis Table 3.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 3.7 · source-reported

Design of the continuation experiments

Experiment          Question                       Approach

EQ1: Grok cita-     Why did Grok’s accu-           40 Australian and United Kingdom
tion sensitivity    racy decline when cita-        cases retested under five formats:
                    tion context was added?        No Clue, With Clue, Late Reveal,
                                                   Partial Clue and Jurisdiction Only.
EQ2:                Did the 20/20 United           Examination of whether responses
DeepSeek’s          Kingdom No Clue result         reproduced the original case citation
United Kingdom      reflect transferable per-      from sanitised facts, comparison with
result              formance or possible case      other jurisdictions, and a rerun of the 20
                    familiarity?                   United Kingdom cases.
EQ3: ChatGPT        Could            alternative   Selected errors retested under
prompt sensitiv-    prompts        change    se-   Adversarial Pushback, Jurisdiction
ity                 lected    Australian    and    Masking and Confidence Penalty
                    United States errors?          prompts.

Thesis Table 3.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.