CrossLaw/Part I · Benchmark/Decision year and significant results
08 / 26 · Part I · BenchmarkThesis §§4.7–4.8; Tables 4.11–4.12

The year pattern does not establish a recency effect.

Decision year and significant results

Key takeaways

  • The 2017 table entry is marked and excluded from interpretation.
  • Six principal statistically significant results are collected in Table 4.12.
  • Per-year case counts were not reported.

What was tested & why

Decision year and significant results

The lowest interpretable year is 2025 at 61.5%, but 2023 and 2024 are closer to 2015. Reported significance belongs to the specific tests and samples shown, not a general prediction about newer cases.

Evidence boundary

The 2017 value is source-marked without explanation. No recency claim is established.

Source: Thesis §§4.7–4.8; Tables 4.11–4.12 · source-reported unless otherwise noted.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Section 11 · Performance by case yearReport lines 633–654

SECTION 11: ANALYSIS BY CASE YEAR

11.1 Recency Effect – Performance by Year

YearCasesAverage Accuracy (No Clue)Notes
2015US_001, US_002, US_004, US_012, US_014, US_015, US_016, US_017, US_018, US_019, AU_001-008, UK_02078.3%Well-established doctrine
2016AU_009-012, US_003, US_00676.5%Strong performance
2017AU_0130%Universal failure
2018AU_014-016, NG_008, NG_016, NG_020, UK_017, UK_01870.2%Mixed performance
2019AU_017, NG_015, NG_016, UK_01976.7%Moderate
2020AU_018, NG_002, NG_01973.3%Moderate
2021AU_019, NG_013, NG_01766.7%Decline
2022NG_003, NG_012, NG_018, UK_013, US_003, US_02075.0%Strong
2023NG_011, UK_001, UK_003, UK_007, UK_008, UK_009, UK_010, UK_012, US_005, US_008, US_010, US_01178.1%Strong
2024NG_007, NG_014, UK_002, UK_004, UK_005, UK_011, UK_014, UK_015, US_007, AU_02076.7%Moderate (includes 2026 cases)
2025NG_001, NG_004, NG_005, NG_006, NG_009, NG_010, NG_011, US_01361.5%Sharp decline
2026UK_002, UK_00470.8%Mixed (limited data)

11.2 Key Findings on Recency

  • Sharp performance decline for 2025 cases – average accuracy drops to 61.5%, compared to 78.3% for 2015 cases. This 16.8% degradation confirms a significant recency effect.
  • US_013 (2025) most difficult – only 3/6 models correct in No-Clue, 2/6 in With-Clue, demonstrating frontier of training data.
  • 2026 UK cases show mixed results – UK_002 (Hexagon Housing) correctly predicted by 5/6 models, UK_004 (G4S) correctly predicted by 4/6, suggesting some recent cases are well-covered.
  • 2017 outlier (AU_013) – 0% accuracy, demonstrating that case difficulty can override recency effects.
  • Implication for practitioners – exercise heightened caution with cases decided after 2023, as model training data may be incomplete.

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 4.11 · source-reported

Average No Clue accuracy across models by year of decision (%)

                 Year          2015   2016     2017   2018    2019    2020
                 Accuracy      78.3   76.5     0.0∗   70.2    76.7    73.3

                 Year          2021   2022     2023   2024    2025    2026
                 Accuracy      66.7   75.0     78.1   76.7    61.5    70.8
 ∗
  The 2017 value is marked with an asterisk in the source record without further explanation
  and is excluded from interpretation. The number of cases per year is not reported, so each
                                 value may rest on few cases.

Thesis Table 4.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.12 · source-reported

Principal statistically significant results of the primary benchmark

Finding                                 Test                   Result         Effect size

Gemini, United States versus other      Mann–Whitney U         p = 0.0019     Cliff’s δ     =
jurisdictions, No Clue                                                        0.40
ChatGPT, United States versus           Mann–Whitney U         p = 0.0093     Cliff’s δ     =
other jurisdictions, With Clue                                                −0.30
Association between DeepSeek            Chi-square             p = 0.0005     Not reported
correctness and jurisdiction
DeepSeek ordered                        Jonckheere–            p = 0.0067     Not reported
Commonwealth-lineage trend              Terpstra
Gemini No Clue versus With Clue         McNemar                p = 0.0044     Not reported
ChatGPT No Clue versus                  McNemar                p = 0.0347     Not reported
With Clue

Thesis Table 4.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.