Key takeaways
- The 2017 table entry is marked and excluded from interpretation.
- Six principal statistically significant results are collected in Table 4.12.
- Per-year case counts were not reported.
What was tested & why
Decision year and significant results
The lowest interpretable year is 2025 at 61.5%, but 2023 and 2024 are closer to 2015. Reported significance belongs to the specific tests and samples shown, not a general prediction about newer cases.
The 2017 value is source-marked without explanation. No recency claim is established.
Source: Thesis §§4.7–4.8; Tables 4.11–4.12 · source-reported unless otherwise noted.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 11 · Performance by case yearReport lines 633–654
SECTION 11: ANALYSIS BY CASE YEAR
11.1 Recency Effect – Performance by Year
| Year | Cases | Average Accuracy (No Clue) | Notes |
|---|---|---|---|
| 2015 | US_001, US_002, US_004, US_012, US_014, US_015, US_016, US_017, US_018, US_019, AU_001-008, UK_020 | 78.3% | Well-established doctrine |
| 2016 | AU_009-012, US_003, US_006 | 76.5% | Strong performance |
| 2017 | AU_013 | 0% | Universal failure |
| 2018 | AU_014-016, NG_008, NG_016, NG_020, UK_017, UK_018 | 70.2% | Mixed performance |
| 2019 | AU_017, NG_015, NG_016, UK_019 | 76.7% | Moderate |
| 2020 | AU_018, NG_002, NG_019 | 73.3% | Moderate |
| 2021 | AU_019, NG_013, NG_017 | 66.7% | Decline |
| 2022 | NG_003, NG_012, NG_018, UK_013, US_003, US_020 | 75.0% | Strong |
| 2023 | NG_011, UK_001, UK_003, UK_007, UK_008, UK_009, UK_010, UK_012, US_005, US_008, US_010, US_011 | 78.1% | Strong |
| 2024 | NG_007, NG_014, UK_002, UK_004, UK_005, UK_011, UK_014, UK_015, US_007, AU_020 | 76.7% | Moderate (includes 2026 cases) |
| 2025 | NG_001, NG_004, NG_005, NG_006, NG_009, NG_010, NG_011, US_013 | 61.5% | Sharp decline |
| 2026 | UK_002, UK_004 | 70.8% | Mixed (limited data) |
11.2 Key Findings on Recency
- Sharp performance decline for 2025 cases – average accuracy drops to 61.5%, compared to 78.3% for 2015 cases. This 16.8% degradation confirms a significant recency effect.
- US_013 (2025) most difficult – only 3/6 models correct in No-Clue, 2/6 in With-Clue, demonstrating frontier of training data.
- 2026 UK cases show mixed results – UK_002 (Hexagon Housing) correctly predicted by 5/6 models, UK_004 (G4S) correctly predicted by 4/6, suggesting some recent cases are well-covered.
- 2017 outlier (AU_013) – 0% accuracy, demonstrating that case difficulty can override recency effects.
- Implication for practitioners – exercise heightened caution with cases decided after 2023, as model training data may be incomplete.
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Average No Clue accuracy across models by year of decision (%)
Year 2015 2016 2017 2018 2019 2020
Accuracy 78.3 76.5 0.0∗ 70.2 76.7 73.3
Year 2021 2022 2023 2024 2025 2026
Accuracy 66.7 75.0 78.1 76.7 61.5 70.8
∗
The 2017 value is marked with an asterisk in the source record without further explanation
and is excluded from interpretation. The number of cases per year is not reported, so each
value may rest on few cases.Thesis Table 4.11 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Principal statistically significant results of the primary benchmark
Finding Test Result Effect size Gemini, United States versus other Mann–Whitney U p = 0.0019 Cliff’s δ = jurisdictions, No Clue 0.40 ChatGPT, United States versus Mann–Whitney U p = 0.0093 Cliff’s δ = other jurisdictions, With Clue −0.30 Association between DeepSeek Chi-square p = 0.0005 Not reported correctness and jurisdiction DeepSeek ordered Jonckheere– p = 0.0067 Not reported Commonwealth-lineage trend Terpstra Gemini No Clue versus With Clue McNemar p = 0.0044 Not reported ChatGPT No Clue versus McNemar p = 0.0347 Not reported With Clue
Thesis Table 4.12 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.