Key takeaways
- Silent Failure is the study’s label for an incorrect verdict at >50% stated confidence.
- The count fell from 134 No Clue to 88 With Clue.
- Australia had the highest jurisdiction rate, 30.8%.
What was tested & why
Confidence and Silent Failure
Accuracy and confidence reliability are separable. Expected Calibration Error and Brier score are shown in the thesis tables; high confidence is not an independent check of correctness.
Silent Failure is a study-specific operational measure, not a general industry classification.
Source: Thesis §§4.4–4.6; Tables 4.5–4.10 · source-reported unless otherwise noted.
Results / visual evidence
Silent Failure by jurisdiction
Thesis Table 4.8 · 240 predictions per jurisdiction. Percentages are source-reported; approximate counts are in the table.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
Silent Failures by model and jurisdiction (both conditions, 40 predictions per cell)
| Model | Nigeria | Australia | US | UK | Total /160 |
|---|---|---|---|---|---|
| ChatGPT | 6 | 14 | 13 | 6 | 39 |
| Gemini | 9 | 7 | 0 | 7 | 23 |
| Claude | 9 | 11 | 7 | 6 | 33 |
| Grok | 7 | 12 | 2 | 6 | 27 |
| DeepSeek | 17 | 20 | 19 | 3 | 59 |
| Perplexity | 9 | 10 | 13 | 9 | 41 |
| All models | 57 | 74 | 54 | 37 | 222 |
Report-only cell counts. Cross-check: column totals equal thesis Table 4.8 (57, 74, 54, 37), the grand total equals Table 4.7 (222), and the DeepSeek, Perplexity and ChatGPT totals equal Table 4.7’s three highest model counts. The thesis does not publish the individual cells. DeepSeek’s burden was concentrated outside the UK; Gemini had none on US cases.
Silent Failures by model and prompting condition (80 predictions per cell)
| Model | No Clue | With Clue | Change | Incorrect NC | Incorrect WC |
|---|---|---|---|---|---|
| ChatGPT | 25 | 14 | -11 | 25 | 14 |
| Gemini | 18 | 5 | -13 | 18 | 5 |
| Claude | 21 | 12 | -9 | 21 | 12 |
| Grok | 13 | 14 | 1 | 13 | 16 |
| DeepSeek | 32 | 27 | -5 | 32 | 27 |
| Perplexity | 25 | 16 | -9 | 25 | 16 |
| All models | 134 | 88 | -46 | 134 | 90 |
Silent Failure counts are report-only; condition totals (134 and 88, a 34.3% reduction) match thesis Table 4.7. Incorrect counts are 80 minus correct from Tables 4.1–4.2: under No Clue every error was a Silent Failure; under With Clue only two Grok errors were not. Grok was the only model with more Silent Failures after the clue.
Silent Failures by jurisdiction and condition (all six models)
| Group | No Clue | With Clue | Reduction | % reduction |
|---|---|---|---|---|
| Nigeria | 39 | 18 | 21 | 53.8 |
| Australia | 43 | 31 | 12 | 27.9 |
| United States | 29 | 25 | 4 | 13.8 |
| United Kingdom | 23 | 14 | 9 | 39.1 |
| Total | 134 | 88 | 46 | 34.3 |
Report-only split; totals match thesis Tables 4.7–4.8. Shown beside the model tables so the jurisdiction view is not read alone.
Calibration by model: Brier score and Expected Calibration Error (lower is better)
| Model | Brier NC | Brier WC | Brier change % | ECE NC | ECE WC |
|---|---|---|---|---|---|
| ChatGPT | 0.239 | 0.156 | -34.7 | 0.1611 | 0.098 |
| Gemini | 0.184 | 0.055 | -70.1 | 0.1459 | 0.0188 |
| Claude | 0.187 | 0.138 | -26.2 | 0.0375 | 0.0959 |
| Grok | 0.133 | 0.145 | 9 | 0.0571 | 0.0256 |
| DeepSeek | 0.327 | 0.298 | -8.9 | 0.2898 | 0.262 |
| Perplexity | 0.237 | 0.164 | -30.8 | 0.1793 | 0.11 |
Thesis values. DeepSeek had the weakest Brier and ECE in both conditions; Gemini had the lowest With Clue Brier and ECE; Grok had the lowest No Clue Brier and was the only model whose Brier score rose with the clue. Change column is not shaded.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 13 · Notable failure casesReport lines 677–702
SECTION 13: NOTABLE FAILURE CASES
13.1 NG_009 – Ethiopian Airlines v Polaris Bank (Nigeria)
| Condition | Correct | Issue |
|---|---|---|
| No Clue | 0/6 | All models reasoned on merits, missed statute-bar entirely |
| With Clue | 2/6 | Only ChatGPT and Gemini corrected |
Lesson: LLMs systematically underperform on procedural/limitation points – a critical risk for practitioners. This demonstrates a structural bias toward substantive merits reasoning over procedural analysis. Even when the correct legal outcome depends entirely on a limitation period, models default to analyzing the substantive merits of the dispute.
13.2 NG_020 – Atiba v Suberu (Nigeria)
| Condition | Correct | Issue |
|---|---|---|
| No Clue | 0/6 | All models predicted borrower's appeal would succeed, assuming a letter superseded a formal deed |
| With Clue | 5/6 | All except DeepSeek corrected with citation |
Lesson: The parol evidence / document hierarchy problem is extremely common in Nigerian property and banking disputes. All six models defaulted to the wrong intuition – that a subsequent letter from the lender modified the formal mortgage deed. This tells practitioners something specific about where not to trust these tools without verification. The universal failure on this case is as significant as NG_009, revealing a systematic weakness in understanding document hierarchy and integration rules.
13.3 AU_013 – Ecosse Property Holdings (Australia)
| Condition | Correct | Issue |
|---|---|---|
| No Clue | 0/6 | All models applied literal textual construction of "payable by the tenant" |
| With Clue | 3/6 | ChatGPT, Gemini, Grok corrected with citation |
Lesson: Purposive commercial construction is a systematic weakness; citation retrieval helps but is not universal. This case represents the only universal failure in the Australian dataset and demonstrates that models default to literal interpretation even when commercial purpose demands a different construction.
13.4 UK_004 – G4S v Lewis-Ranwell (UK)
| Condition | Correct | Issue |
|---|---|---|
| With Clue | 3/6 | Failed by ChatGPT, Grok, DeepSeek |
Lesson: Illegality defence application to insanity verdicts is genuinely difficult – even with citation. This 2026 judgment tests training data frontiers and shows that even when citations are provided, models may struggle with novel or complex policy questions.
13.5 US_013 – Medical Marijuana v Horn (US)
| Condition | Correct | Issue |
|---|---|---|
| No Clue | 3/6 | Only 3 models correct |
| With Clue | 2/6 | Only Gemini and Grok correct |
Lesson: Most recent case (2025) tests training data frontier – significant recency effect. Even with citation, 4 of 6 models failed, demonstrating that models cannot be relied upon for very recent jurisprudence.
Report Appendix A · Supplementary statistical analyses (logistic regression, ECE, domain chi-square, power)Report lines 2455–2592
APPENDIX A: Supplementary Statistical Analyses Report (Optional)
1. Introduction and Rationale
The primary statistical validation (Part 2 of the main report) established key findings using non parametric tests (Friedman, McNemar, Mann Whitney U, Jonckheere Terpstra, Chi square, Cochran’s Q, coefficient of variation, and reliability diagrams). Those tests were appropriate for the experimental design and provided robust evidence for the headline conclusions.
However, to further strengthen the research and address potential peer review questions, three supplementary statistical evaluations were conducted:
- Mixed effects logistic regression – to quantify the simultaneous effect of Condition, Model, and Legal Domain on accuracy while controlling for case level variance (a main effects logit model was used as an approximation).
- Expected Calibration Error (ECE) – to provide a single numeric measure of confidence calibration for each model and condition, complementing the visual reliability diagrams.
- Domain specific Chi square tests – to test whether accuracy varies significantly across legal domains overall and per jurisdiction.
- Power analysis for non significant results – to determine if the sample size was adequate to detect meaningful effects for the non significant findings (e.g., Grok’s decline with clues, Claude’s marginal improvement).
These analyses are optional but add rigour, confirm the adequacy of the sample size, quantify calibration, and demonstrate the influence of legal domain. They are placed in this Supplementary Statistical Analyses section, which can be appended to Part 2 of the main report as Appendix A.
2. Methodology
All analyses were performed using the complete dataset of 960 verified experiments (20 cases × 4 jurisdictions × 6 models × 2 conditions). Data preparation followed the same protocol as the main study (sanitised facts, binary verdict scores, confidence scores scaled to [0,1]). The code was executed in a Google Colab environment using Python with libraries statsmodels, scikit-learn, pingouin, and numpy.
2.1 Logistic Regression (Main Effects)
A logistic regression model was fitted with the binary outcome correct (1 = correct verdict, 0 = incorrect). Fixed effects included:
- Condition_num (0 = No Clue, 1 = With Clue)
- Model (categorical, reference = ChatGPT)
- Legal_Domain (categorical, reference = not explicitly set; the output shows comparisons for many domain levels)
The model was estimated using maximum likelihood (method='bfgs') with 1000 iterations to ensure convergence. No random effects were included due to the complexity of crossed random effects in Python; the model serves as an approximation of the main fixed effects. Consequently, within case correlation is not accounted for, but the fixed effect coefficients are consistent with the non parametric tests.
2.2 Expected Calibration Error (ECE)
ECE quantifies the average absolute difference between predicted confidence and observed accuracy across bins. A lower ECE indicates better calibration.
For each model–condition pair, predicted confidences were divided into 10 equal width bins [0,0.1), [0.1,0.2), … , [0.9,1.0]. For each bin, the mean confidence and observed accuracy were calculated. ECE is the weighted average of the absolute differences:
"ECE"=∑_(b=1)^B▒n_b/N∣〖"acc" 〗_b-〖"conf" 〗_b∣
where n_b is the number of predictions in bin b, N total predictions, 〖"acc" 〗_b observed accuracy, and 〖"conf" 〗_b mean confidence. Empty bins were safely ignored.
2.3 Domain specific Chi square Tests
For each jurisdiction and for all jurisdictions combined, a contingency table of Legal_Domain vs. correct was constructed. A chi square test of independence was performed to determine whether the distribution of correct/incorrect answers differs significantly by legal domain. A significant result indicates that some legal domains are systematically easier or harder for LLMs.
2.4 Power Analysis for Non significant Results
Two non significant results from the main McNemar tests were examined:
- Grok’s decline with clues (−2.5%, p = 0.8238)
- Claude’s improvement (+11.25%, p = 0.0784)
For each, the effect size Cohen’s h was calculated:
h=2arcsin(√(p_1 ))-2arcsin(√(p_2 ))
where p_1 and p_2 are the proportions correct in the two conditions (No Clue vs With Clue). Using proportion_effectsize and tt_ind_solve_power (two sample independent proportions power) with n=80 per group, α=0.05, two tailed, the statistical power to detect the observed effect was estimated. This provides a check on whether non significance could be due to insufficient sample size.
Note: The independent samples approximation is conservative; the exact McNemar power depends on discordant pairs, but the directional conclusions remain valid.
3. Results
3.1 Logistic Regression (Main Effects)
The logistic regression converged successfully (Pseudo R squared = 0.2089, log likelihood ratio test p < 0.0001). The full coefficient table is presented below.
Table 1: Logistic Regression Results (Main Effects Only)
| Variable | Coefficient | Std. Error | z | P>|z| | 95% CI |
|----------|------------|------------|----|-------|--------|
| Intercept | 2.0927 | 1.081 | 1.935 | 0.053 | [-0.027, 4.212] |
| Condition_num | 0.6729 | 0.176 | 3.827 | 0.000 | [0.328, 1.018] |
| Model (ref: ChatGPT) | | | | | |
| Claude | 0.2678 | 0.299 | 0.895 | 0.371 | [-0.319, 0.854] |
| DeepSeek | -0.7591 | 0.279 | -2.719 | 0.007 | [-1.306, -0.212] |
| Gemini | 0.7973 | 0.322 | 2.477 | 0.013 | [0.167, 1.428] |
| Grok | 0.5159 | 0.309 | 1.672 | 0.095 | [-0.089, 1.121] |
| Perplexity | -0.0835 | 0.289 | -0.289 | 0.773 | [-0.650, 0.483] |
| Legal_Domain (selected significant) | | | | | |
| Commercial Law (Banking) | -4.2719 | 1.327 | -3.219 | 0.001 | [-6.873, -1.671] |
| Commercial Law (Ethics) | -3.3053 | 1.239 | -2.667 | 0.008 | [-5.734, -0.876] |
| Contract Law (Construction) | -3.3053 | 1.239 | -2.667 | 0.008 | [-5.734, -0.876] |
| Contract Law (Mortgage) | -2.9211 | 1.226 | -2.383 | 0.017 | [-5.324, -0.518] |
| Equity (Trusts) | -2.5565 | 1.222 | -2.092 | 0.036 | [-4.951, -0.161] |
| Civil Law (RICO) | -2.5565 | 1.222 | -2.092 | 0.036 | [-4.951, -0.161] |
| Intellectual Property (Trademark) | -2.0661 | 1.121 | -1.842 | 0.065 | [-4.264, 0.132] |
Note: Only a subset of legal domains with p < 0.10 are shown for brevity; the full table is available in the code output.
Key findings:
- The Condition_num coefficient (+0.6729, p < 0.001) confirms that including citations significantly increases the log odds of a correct verdict, consistent with the McNemar test.
- DeepSeek is significantly worse than ChatGPT (p = 0.007), while Gemini is significantly better (p = 0.013). Grok is marginally better (p = 0.095).
- Several legal domains (e.g., Commercial Law (Banking), Contract Law (Construction), Equity (Trusts)) show large negative coefficients, indicating that models perform worse on those domains relative to the reference category.
3.2 Expected Calibration Error (ECE)
Table 2 reports the ECE for each model and condition. Lower ECE indicates better alignment between confidence and observed accuracy.
ECE < 0.05 is excellent, 0.05–0.10 is good, > 0.20 indicates serious miscalibration.
Table 2: Expected Calibration Error by Model and Condition
| Model | Condition | ECE |
|---|---|---|
| Claude | No Clue | 0.0375 |
| Grok | No Clue | 0.0571 |
| Gemini | No Clue | 0.1459 |
| ChatGPT | No Clue | 0.1611 |
| Perplexity | No Clue | 0.1793 |
| DeepSeek | No Clue | 0.2898 |
| Gemini | With Clue | 0.0188 |
| Grok | With Clue | 0.0256 |
| Claude | With Clue | 0.0959 |
| ChatGPT | With Clue | 0.0980 |
| Perplexity | With Clue | 0.1100 |
| DeepSeek | With Clue | 0.2620 |
Interpretation:
- Best calibrated: Claude (No Clue, ECE=0.0375) and Gemini (With Clue, ECE=0.0188) – their confidences reflect actual accuracy almost perfectly.
- Worst calibrated: DeepSeek in both conditions (ECE = 0.290 and 0.262) – it is dangerously overconfident, frequently expressing high confidence on incorrect predictions. This confirms the “Silent Failure” risk identified in the descriptive analysis.
- Grok has low ECE in both conditions, indicating good calibration despite its accuracy decline with clues.
- ChatGPT and Perplexity show moderate miscalibration, improving with clues.
3.3 Domain specific Chi square Tests
Table 3: Chi square Test for Independence of Legal Domain and Correctness
| Jurisdiction | χ² | df | p-value | Significance (α=0.05) |
|---|---|---|---|---|
| Overall (all jurisdictions) | 162.994 | 73 | 0.0000 | Significant |
| Nigeria | 65.388 | 15 | 0.0000 | Significant |
| United States | 25.942 | 14 | 0.0173 | Significant |
| Australia | 45.992 | 16 | 0.0003 | Significant |
| United Kingdom | 22.080 | 18 | 0.2803 | Not significant |
Findings:
- Accuracy varies significantly by legal domain overall and in Nigeria, the United States, and Australia.
- In the United Kingdom, domain was not a significant factor (p = 0.2803). This suggests that, uniquely among the four jurisdictions, UK legal domains are handled more uniformly by the models. This may reflect the high quality centralised reporting of UK case law (BAILII) and/or DeepSeek’s memorisation of UK cases, which reduces domain specific variation.
3.4 Power Analysis for Non significant Results
Table 4: Power to Detect Observed Effect Sizes (Independent proportions approximation)
| Model | Comparison | p value (McNemar) | p₁ (NC) | p₂ (WC) | Cohen’s h | Effect size | Power (two sample independent) |
|---|---|---|---|---|---|---|---|
| Grok | NC vs WC | 0.8238 | 0.8375 | 0.8000 | 0.097 | Negligible | 0.094 |
| Claude | NC vs WC | 0.0784 | 0.7375 | 0.8500 | –0.280 | Small medium | 0.422 |
Interpretation:
- Grok: The observed decline (Cohen’s h = 0.097) is negligible. The power to detect such a small effect with 80 cases is only 9.4%, meaning the non significant result is not due to low power but because the true effect is practically zero. The descriptive finding that Grok declines with clues is not statistically meaningful.
- Claude: The improvement (Cohen’s h = 0.280) is a small to medium effect. The power of 42% indicates that the current sample size had only a moderate chance of detecting this effect. Hence the marginal p value (0.0784) could be a false negative – the improvement may be real but requires a larger sample to confirm. This supports the descriptive observation that Claude’s improvement (+11.25%) is consistent across jurisdictions and may be genuine.
- Overall sample size (80 cases per condition) is adequate to detect a medium effect (Cohen’s h ≥ 0.3) with power > 0.8 (as shown in the output). Non significant results with smaller effects are therefore not due to insufficient data.
Note: These power calculations use an independent samples proportion test, a conservative approximation for the paired McNemar design. Actual power for McNemar would be slightly higher, but the directional conclusions remain valid.
4. Integration with Prior Findings
4.1 Consistency with Descriptive and Main Statistical Analysis
- Condition effect (+9.1% average improvement) is strongly confirmed by the logistic regression (p < 0.001) and the low ECE of Gemini with clues.
- DeepSeek’s poor performance is reinforced: logistic regression shows it is significantly worse than ChatGPT (p = 0.007), and its high ECE (>0.26) confirms overconfidence, aligning with the “Silent Failure” warning in the descriptive section.
- Gemini’s excellence with clues is supported by the lowest ECE (0.0188) and the significant positive coefficient (p = 0.013).
- Grok’s clue induced decline is shown to be negligible (h = 0.097) and not statistically meaningful; however, its calibration remains good (ECE = 0.0256 with clues). This clarifies that while Grok’s accuracy does not improve, it does not become dangerously miscalibrated.
- Domain variability was already observed descriptively (e.g., NG_020 and AU_013 domain failures). The chi square tests quantify this: accuracy differs by domain in three of four jurisdictions, and the UK is the exception.
4.2 Contributions of Supplementary Analyses
| Supplementary Test | Added Value |
|---|---|
| Logistic regression | Quantifies the simultaneous effect of Condition, Model, and Domain, confirming the strength of citation context and identifying domains of systematic difficulty. |
| Expected Calibration Error | Provides a single numeric calibration metric, objectively showing that DeepSeek is dangerously overconfident, while Gemini and Claude are well calibrated. |
| Domain chi square | Empirically demonstrates that legal domain significantly impacts performance, with the UK being uniquely uniform – a new finding that may reflect data quality. |
| Power analysis | Reassures that non significant results (Grok’s decline) are truly negligible, while Claude’s marginal improvement may warrant further study. |
4.3 Implications for the Industrial Translation
- ECE reinforces the “Silent Failure” warning for DeepSeek: not only is its accuracy low, but its confidence is systematically misleading. The practitioner’s Duty of Inquiry Checklist must mandate independent verification of DeepSeek outputs regardless of confidence.
- Domain chi square identifies specific legal domains (e.g., Commercial Law (Banking), Contract Law (Construction), Equity (Trusts)) as high risk areas where extra caution is required.
- Power analysis provides evidence that the sample size (80 cases × 2 conditions) is sufficient for medium sized effects, strengthening the credibility of the null findings (e.g., Grok’s decline).
5. Placement in the Overall Report
These supplementary analyses are not required for the main conclusions but provide valuable depth and address potential methodological questions from reviewers. They should be placed as:
6. Conclusion
The supplementary statistical analyses confirm and extend the primary findings:
- Citation context significantly improves accuracy (logistic regression p < 0.001).
- DeepSeek is both low accuracy and severely overconfident (ECE > 0.26), validating the “Silent Failure” risk.
- Gemini and Claude are well calibrated, with Gemini achieving near perfect calibration with clues (ECE = 0.0188).
- Legal domain affects performance in Nigeria, US, and Australia, but not in the UK – a novel observation that may reflect training data characteristics.
- The sample size (80 cases) is sufficient to detect medium effects; Grok’s non significant decline is truly negligible, while Claude’s marginal improvement may merit further investigation.
These optional analyses add rigour, address potential critique, and provide deeper insight into model behaviour. They are now ready for inclusion in the final thesis or paper.
SILENT FAILURE ANALYSIS AND FINDINGS — CROSSLAW BENCHMARK
A Comprehensive Analysis of Confidently Wrong LLM Predictions Across Four Common Law Jurisdictions
Silent Failure analysis and findings reportReport lines 2593–2907
Prepared for: HUMN4002 RD2 Viva Voce | Western Sydney University
Date: Autumn 2026
Analytical Framework: Silent Failure = EV < -0.5 (Incorrect Verdict + High Confidence)
Data Source: 960 verified interactions (80 cases × 6 models × 2 conditions)
EXECUTIVE SUMMARY
Silent Failure — confident wrongness without warning — represents a significant risk in legal AI deployment. A model that is wrong but expresses high confidence creates outputs that appear authoritative and persuasive to practitioners, making errors difficult to detect and correct.
Across 960 controlled interactions, we identified 222 Silent Failures, representing 23.1% of all predictions. Citation context reduced Silent Failures by 34.3% (134 → 88), confirming that prompt design materially improves model reliability.
Australia recorded the highest Silent Failure burden within the benchmark (74, 30.8%), followed by Nigeria (57, 23.8%), United States (54, 22.5%), and United Kingdom (37, 15.4%). DeepSeek accounted for the largest model-specific burden (59 Silent Failures), with its highest observed burdens in Australia (20), the United States (19), and Nigeria (17). Gemini demonstrated the strongest Silent Failure safety profile among the models tested, producing only 23 Silent Failures total and zero in the United States.
Brier Score confirms the EV findings: DeepSeek had the poorest calibration quality observed (Brier = 0.2978 With Clue) and the highest Silent Failure burden (59), while Gemini achieved the strongest With Clue reliability profile (Brier = 0.0552, Silent Failures = 23).
Critical cases in Nigeria and the United States produced EV = -0.95 to -1.0, demonstrating that models can be perfectly confident while being completely wrong — a pattern that would remain invisible under conventional accuracy based benchmarking.
1. DEFINITION AND METHODOLOGY
1.1 Silent Failure Definition
Accuracy alone does not capture the practical risks associated with deploying LLMs in legal contexts. A model may produce an incorrect answer while expressing very high confidence, creating a particularly dangerous form of error that can appear persuasive to practitioners.
Following the benchmark framework, a Silent Failure was defined as any prediction with an Expected Value (EV) below −0.5.
| Condition | Threshold |
|---|---|
| Verdict Score | 0 (incorrect) |
| Expected Value (EV) | < -0.5 (confidently wrong) |
1.2 Expected Value Formula
The Expected Value metric combines accuracy and confidence into a single risk adjusted measure:
| Outcome | Formula |
|---|---|
| Correct Prediction | EV = +Confidence / 100 |
| Incorrect Prediction | EV = −Confidence / 100 |
Interpretation:
| EV Range | Meaning |
|---|---|
| EV > 0.5 | Reliable prediction (correct with moderate confidence) |
| 0 < EV < 0.5 | Correct but uncertain |
| -0.5 < EV < 0 | Incorrect but appropriately cautious |
| EV < -0.5 | ⚠️ SILENT FAILURE — Confidently wrong; worse than random guessing |
1.3 Brier Score Definition
Brier Score measures the mean squared difference between predicted probabilities (Confidence/100) and actual outcomes (0 or 1). Lower Brier Scores indicate better calibration.
Note on Brier Score Interpretation: The thresholds used in this report (e.g., 0.0–0.1 as "excellent") are heuristic guidelines for comparative analysis within this benchmark rather than universally accepted calibration thresholds for legal AI systems. They are intended to facilitate model comparison rather than establish absolute standards.
| Brier Score Range | Heuristic Interpretation |
|---|---|
| 0.0 – 0.1 | Excellent (relative to benchmark) |
| 0.1 – 0.2 | Good (relative to benchmark) |
| 0.2 – 0.3 | Moderate (relative to benchmark) |
| > 0.3 | Poor (relative to benchmark) |
1.4 Dataset Overview
| Parameter | Value |
|---|---|
| Total Cases | 80 |
| Jurisdictions | Nigeria, Australia, United Kingdom, United States |
| Models | ChatGPT 4.1, Gemini 1.5 Pro, Claude 3 Opus, Grok 2, DeepSeek V3, Perplexity Sonar |
| Prompt Conditions | No Clue (facts only), With Clue (facts + citation + jurisdiction + year + domain) |
| Total Interactions | 960 (80 total cases × 6 models × 2 conditions) |
| Silent Failure Definition | EV < -0.5 |
2. OVERALL SILENT FAILURE BURDEN
Across the 960 benchmark interactions, 222 predictions met the Silent Failure threshold, representing 23.1% of all model outputs.
This finding indicates that almost one in four legal AI predictions within the benchmark were confidently incorrect enough to satisfy the Silent Failure criterion. The result demonstrates that accuracy alone substantially understates deployment risk.
3. SILENT FAILURE BURDEN ACROSS JURISDICTIONS
Table 1. Silent Failures by Jurisdiction
| Jurisdiction | Silent Failures | Total Predictions | Rate |
|---|---|---|---|
| Australia | 74 | 240 | 30.8% |
| Nigeria | 57 | 240 | 23.8% |
| United States | 54 | 240 | 22.5% |
| United Kingdom | 37 | 240 | 15.4% |
| TOTAL | 222 | 960 | 23.1% |
Key Observations
- Australia recorded the highest Silent Failure rate within the benchmark (30.8%) — almost one in three predictions were confidently wrong. This suggests that the evaluated models were less reliable on Australian appellate law tasks than on the other jurisdictions examined.
- Nigeria recorded the second highest rate (23.8%) — nearly one in four predictions were Silent Failures. This confirms that low resource jurisdictions face compounding risks.
- The United Kingdom recorded the lowest Silent Failure rate (15.4%) — reflecting better training data coverage and model calibration for UK law.
- The UK was the only jurisdiction below 20% — suggesting that high resource, well digitised legal systems benefit from more reliable AI calibration.
4. SILENT FAILURE DISTRIBUTION BY MODEL
Table 2. Silent Failure Counts by Model
| Model | Silent Failures | % of Total |
|---|---|---|
| DeepSeek | 59 | 26.6% |
| Perplexity | 41 | 18.5% |
| ChatGPT | 39 | 17.6% |
| Claude | 33 | 14.9% |
| Grok | 27 | 12.2% |
| Gemini | 23 | 10.4% |
| TOTAL | 222 | 100% |
Key Observations
- DeepSeek accounted for 26.6% of all Silent Failures (59 of 222) — DeepSeek was the highest risk model in terms of calibration failure within this benchmark.
- Gemini had the lowest observed Silent Failure burden (23) — Gemini was the lowest risk model among those tested.
- Grok had 27 Silent Failures despite high accuracy (83.8%) — Grok was accurate when correct but showed overconfidence when wrong.
- Perplexity's high Silent Failure count (41) — retrieval first models may be more prone to overconfidence when retrieval is incorrect.
5. SILENT FAILURE HEATMAP (MODEL × JURISDICTION)
Table 3. Silent Failure Counts by Model and Jurisdiction
| Model | Nigeria | Australia | United States | United Kingdom | Total |
|---|---|---|---|---|---|
| ChatGPT | 6 | 14 | 13 | 6 | 39 |
| Gemini | 9 | 7 | 0 | 7 | 23 |
| Claude | 9 | 11 | 7 | 6 | 33 |
| Grok | 7 | 12 | 2 | 6 | 27 |
| DeepSeek | 17 | 20 | 19 | 3 | 59 |
| Perplexity | 9 | 10 | 13 | 9 | 41 |
| Total | 57 | 74 | 54 | 37 | 222 |
Key Observations
| Finding | Insight |
|---|---|
| DeepSeek in Australia had the highest observed cell burden (20) | Australian practitioners face elevated risk when using DeepSeek |
| DeepSeek in the United States was second highest (19) | US practitioners also face elevated Silent Failure risk with DeepSeek |
| DeepSeek in Nigeria was also high (17) | Nigerian practitioners face serious risk with DeepSeek deployment |
| Gemini in the US had zero Silent Failures (0) | Gemini was exceptionally well calibrated for US law within this benchmark |
| Gemini had the lowest total Silent Failures (23) | Gemini was the lowest risk model overall |
| DeepSeek in the UK had the lowest cell burden (3) | DeepSeek's UK specialised knowledge was well calibrated, but this does not indicate safety elsewhere |
6. IMPACT OF CITATION CONTEXT ON SILENT FAILURE RISK
One of the most important findings concerns the effect of citation context.
Table 4. Silent Failures by Condition and Jurisdiction
| Jurisdiction | No Clue | With Clue | Reduction | % Reduction |
|---|---|---|---|---|
| Nigeria | 39 | 18 | 21 | 53.8% |
| Australia | 43 | 31 | 12 | 27.9% |
| United States | 29 | 25 | 4 | 13.8% |
| United Kingdom | 23 | 14 | 9 | 39.1% |
| TOTAL | 134 | 88 | 46 | 34.3% |
Key Observations
- Citation context reduced Silent Failures by 34.3% overall (134 → 88). This is one of the strongest findings: prompt design materially improved model reliability.
- Nigeria showed the largest absolute reduction (39 → 18, -53.8%) — citation clues most helped Nigerian legal reasoning, where models otherwise struggled.
- The United Kingdom showed strong reduction (23 → 14, -39.1%) — citation context significantly improved UK law calibration.
- The United States showed the smallest reduction (29 → 25, -13.8%) — clues did not substantially reduce Silent Failures in US law, possibly because models were already overconfident in US legal contexts.
- Australia showed moderate improvement (43 → 31, -27.9%) — clues helped but did not fully resolve Australian calibration issues.
7. SILENT FAILURE BY MODEL × CONDITION
Table 5. Model Specific Silent Failures by Condition
| Model | No Clue | With Clue | Difference | % Reduction |
|---|---|---|---|---|
| ChatGPT | 25 | 14 | -11 | 44.0% |
| Gemini | 18 | 5 | -13 | 72.2% |
| Claude | 21 | 12 | -9 | 42.9% |
| Grok | 13 | 14 | +1 | -7.7% (worsens) |
| DeepSeek | 32 | 27 | -5 | 15.6% |
| Perplexity | 25 | 16 | -9 | 36.0% |
| TOTAL | 134 | 88 | -46 | 34.3% |
Key Observations
- Every model except Grok improved with citation context — Grok showed a slight increase in Silent Failures with clues (+1), consistent with its accuracy decline.
- Gemini showed the largest improvement (-13, 72.2%) — citation context dramatically improved Gemini's calibration.
- DeepSeek showed the smallest improvement (-5, 15.6%) — DeepSeek remained dangerously overconfident even with clues.
- Gemini had the lowest Silent Failure count in both conditions (18 No Clue, 5 With Clue) — lowest risk model overall.
8. BRIER SCORE CALIBRATION FINDINGS
Brier Score measures the mean squared difference between predicted probabilities (Confidence/100) and actual outcomes (0 or 1). Lower Brier Scores indicate better calibration.
Table 6. Average Brier Score by Model and Condition
| Model | No Clue Brier | With Clue Brier | Interpretation |
|---|---|---|---|
| ChatGPT | 0.2387 | 0.1559 | Improves with clues |
| Gemini | 0.1840 | 0.0552 | Best With Clue |
| Claude | 0.1870 | 0.1375 | Improves with clues |
| Grok | 0.1328 | 0.1447 | Best No Clue; slight decline with clues |
| DeepSeek | 0.3274 | 0.2978 | Worst in both conditions |
| Perplexity | 0.2368 | 0.1642 | Improves with clues |
Key Observations
- Gemini achieved the best With Clue calibration observed (Brier = 0.0552) — exceptionally well calibrated when citation context was provided.
- Grok achieved the best No Clue calibration observed (Brier = 0.1328) — well calibrated without clues, consistent with its stable performance.
- DeepSeek had the poorest calibration observed in both conditions — Brier > 0.29 in both No Clue and With Clue, confirming dangerous overconfidence.
- All models except Grok showed improved calibration with citation context — Grok's Brier slightly worsened (0.1328 → 0.1447), consistent with its accuracy decline.
9. RELATIONSHIP BETWEEN BRIER SCORE, EXPECTED VALUE, AND SILENT FAILURE RISK
Although Brier Score and Silent Failure measure different aspects of model behaviour, the two metrics were broadly aligned in this benchmark. A clear association was observed between poorer calibration (higher Brier Scores) and higher Silent Failure counts. Models with poorer calibration generally produced more Silent Failures, suggesting that calibration quality appears to be a useful indicator of deployment risk within this benchmark.
Table 7. Brier Score vs Silent Failures by Model
| Model | With Clue Brier | Silent Failures | Interpretation |
|---|---|---|---|
| DeepSeek | 0.2978 | 59 | Poorest calibration + highest Silent Failures |
| Perplexity | 0.1642 | 41 | Moderate calibration |
| ChatGPT | 0.1559 | 39 | Good calibration |
| Claude | 0.1375 | 33 | Good calibration |
| Grok | 0.1447 | 27 | Moderate calibration (slightly worsens with clues) |
| Gemini | 0.0552 | 23 | Best calibration + lowest Silent Failures |
Key Observations
- A clear association is observed between poorer calibration (higher Brier Scores) and higher Silent Failure counts.
- DeepSeek showed the poorest calibration observed (Brier = 0.2978) and the highest Silent Failure burden (59).
- Gemini showed the best calibration observed (Brier = 0.0552) and the lowest Silent Failure burden (23).
- Grok showed good No Clue calibration but slightly worsened with clues, consistent with its unique decline pattern.
- Brier Score confirms the EV findings — the same models that were overconfident (high Brier) were also the highest risk in terms of Silent Failure burden.
10. CRITICAL SILENT FAILURE CASES
Several cases demonstrated extreme Silent Failure behaviour.
Table 8. Example Silent Failure Cases (EV < -0.8)
| Case | Case Title | Jurisdiction | Model | Condition | Verdict | Confidence | EV |
|---|---|---|---|---|---|---|---|
| NG_004 | F.H.A. v. Oyedeji | Nigeria | DeepSeek | No Clue | 0 | 95% | -0.95 |
| NG_009 | Ethiopian Airlines v. Polaris Bank | Nigeria | DeepSeek | No Clue | 0 | 95% | -0.95 |
| NG_020 | Atiba Iyalamu v. Suberu | Nigeria | DeepSeek | No Clue | 0 | 95% | -0.95 |
| US_008 | Amgen Inc. v. Sanofi | United States | DeepSeek | No Clue | 0 | 100% | -1.00 |
| US_010 | Bartenwerfer v. Buckley | United States | DeepSeek | No Clue | 0 | 100% | -1.00 |
| UK_004 | G4S Health Services v. Lewis Ranwell | United Kingdom | DeepSeek | With Clue | 0 | 100% | -1.00 |
In each case, the model produced an incorrect legal prediction while maintaining near maximal or perfect confidence. These outputs would likely appear persuasive to practitioners despite being wrong.
Key Observations:
- DeepSeek in Nigeria produced EV = -0.95 on multiple cases — near maximal confidence on incorrect predictions.
- DeepSeek in the United States produced EV = -1.00 on two cases — 100% confidence on wrong answers.
- DeepSeek in the United Kingdom produced EV = -1.00 on one With_Clue case — citation context did not prevent the Silent Failure.
Such cases illustrate why Silent Failure analysis provides information that traditional accuracy metrics cannot capture.
11. VISUALISATION READY SUMMARY
Silent Failure Risk by Jurisdiction
text
Australia ████████████████████████████████████████████ 30.8% (74/240)
Nigeria ██████████████████████████████████████████ 23.8% (57/240)
United States ████████████████████████████████████████ 22.5% (54/240)
United Kingdom █████████████████████████████████████████ 15.4% (37/240)
Silent Failure Burden by Model
text
DeepSeek ████████████████████████████████████████████ 59
Perplexity ████████████████████████████████████████████ 41
ChatGPT ████████████████████████████████████████████ 39
Claude ████████████████████████████████████████████ 33
Grok ████████████████████████████████████████████ 27
Gemini ████████████████████████████████████████████ 23
Condition Effect
text
No Clue ████████████████████████████████████████████████████████████████████████████████████ 134
With Clue ████████████████████████████████████████████████████████████████████████████████████ 88
↓ 34.3% reduction (46 fewer Silent Failures)
Brier Score by Model (With Clue)
text
DeepSeek ████████████████████████████████████████████ 0.2978
Grok ████████████████████████████████████████████ 0.1447
Perplexity ████████████████████████████████████████████ 0.1642
ChatGPT ████████████████████████████████████████████ 0.1559
Claude ████████████████████████████████████████████ 0.1375
Gemini ████████████████████████████████████████████ 0.0552
12. PRACTITIONER IMPLICATIONS
12.1 Model Selection Guidance
| Use Case | Recommended Model | Rationale |
|---|---|---|
| Lowest observed Silent Failure Risk | Gemini | Only 23 Silent Failures total; 0 in US; Brier = 0.0552 |
| Highest observed Silent Failure Risk | DeepSeek | 59 Silent Failures; poorest calibration (Brier = 0.2978) |
| Australian Law | Exercise caution with DeepSeek | 20 Silent Failures in Australia — highest cell burden |
| Nigerian Law | Exercise caution with DeepSeek | 17 Silent Failures in Nigeria — critically high |
| US Law | Gemini | No Silent Failures observed in the US benchmark |
| UK Law | Gemini or Claude | Low Silent Failure rates; Gemini (7), Claude (6) |
12.2 Prompt Design Guidance
| Insight | Implication |
|---|---|
| Citation context reduces Silent Failures by 34.3% | Always include case citation, jurisdiction, year, and legal domain |
| Nigeria shows largest improvement (-53.8%) | Nigerian practitioners benefit most from citation anchoring |
| US shows smallest improvement (-13.8%) | US practitioners should not rely solely on citations to fix calibration |
| DeepSeek improves only 15.6% with clues | DeepSeek remains high risk even with optimal prompting |
12.3 Brier Score and Calibration Guidance
| Insight | Implication |
|---|---|
| Gemini Brier = 0.0552 (With Clue) | Gemini's confidence levels are well calibrated when citations are provided |
| DeepSeek Brier = 0.2978 (With Clue) | DeepSeek's confidence is unreliable; always verify independently |
| Grok Brier = 0.1328 (No Clue) | Grok is well calibrated without clues; confidence is more trustworthy |
| Brier Score association with Silent Failures | Use Brier Score as a proxy for deployment safety |
12.4 DeepSeek Risk Warning
⚠️ Exercise caution when deploying DeepSeek in Australian, Nigerian, and United States legal contexts and require independent verification of outputs.
Evidence:
| Metric | Value |
|---|---|
| Total Silent Failures | 59 — highest of any model |
| Silent Failures in Australia | 20 — highest cell burden |
| Silent Failures in the United States | 19 — second highest cell burden |
| Silent Failures in Nigeria | 17 — critically high |
| With Clue Improvement | Only 15.6% — calibration remained dangerously poor |
| With Clue Brier Score | 0.2978 — poorest calibration observed |
| Extreme Cases | US_008 (EV = -1.00), US_010 (EV = -1.00), UK_004 (EV = -1.00) |
13. THEORETICAL CONTRIBUTIONS
13.1 Silent Failure Quantification
This study provides a systematic quantification of Silent Failures across multiple common law jurisdictions. Prior work identified Silent Failures conceptually (Pathak et al., 2025; Potts & Sudhof, 2026), but did not measure them across jurisdictions, models, and prompt conditions in a controlled benchmark.
13.2 Jurisdictional Risk Mapping
The finding that Australia recorded the highest Silent Failure rate (30.8%) while the UK recorded the lowest (15.4%) confirms that Silent Failure risk is jurisdiction dependent within this benchmark. This challenges the assumption that Silent Failure is a uniform model property.
13.3 Prompt Design as Risk Mitigation
The finding that citation context reduces Silent Failures by 34.3% confirms that prompt design materially improves calibration. This is one of the strongest practical findings: simple changes to input structure reduce the most dangerous failure mode.
13.4 Model Specific Risk Profiles
The identification of DeepSeek as the highest Silent Failure model and Gemini as the lowest provides a model specific Silent Failure risk taxonomy for legal AI within this benchmark framework.
13.5 Brier Score as an Indicator of Silent Failure Risk
The clear association observed between Brier Score and Silent Failure burden suggests that calibration quality may be a useful indicator of deployment safety. Models with poorer calibration (high Brier) consistently produced more Silent Failures within this benchmark.
13.6 Extreme Cases as Evidence of Overconfidence Risk
The existence of predictions with EV = -1.0 (100% confidence on incorrect verdicts) is consistent with concerns raised in the literature regarding overconfident language model outputs (Bender et al., 2021). These cases demonstrate that perfect confidence does not imply correct reasoning.
14. SUMMARY OF KEY FINDINGS
| # | Finding | Value | Implication |
|---|---|---|---|
| 1 | Total Silent Failures | 222 (23.1%) | Silent Failure is pervasive within the benchmark |
| 2 | Highest Silent Failure Jurisdiction | Australia (74, 30.8%) | Australian appellate law tasks showed highest risk |
| 3 | Second Highest | Nigeria (57, 23.8%) | Low resource jurisdictions are vulnerable |
| 4 | Lowest | United Kingdom (37, 15.4%) | High resource systems showed lower risk |
| 5 | Highest Silent Failure Model | DeepSeek (59) | DeepSeek showed highest observed risk |
| 6 | Lowest Silent Failure Model | Gemini (23) | Gemini showed lowest observed risk |
| 7 | Worst Cell | DeepSeek in Australia (20) | Highest risk combination identified |
| 8 | Second Worst Cell | DeepSeek in US (19) | US practitioners also face elevated risk |
| 9 | Best Cell | Gemini in US (0) | No Silent Failures observed in the US benchmark |
| 10 | No Clue Silent Failures | 134 | Overconfidence is worse without context |
| 11 | With Clue Silent Failures | 88 | Citation context reduces risk |
| 12 | Silent Failure Reduction | 46 (34.3%) | Prompt design materially improves safety |
| 13 | DeepSeek Extreme Cases | EV = -1.0 (US_008, US_010) | Perfect confidence on wrong answers |
| 14 | Nigerian Critical Cases | EV = -0.95 (NG_004, NG_009, NG_020) | Near maximal confidence on wrong answers |
| 15 | Best With Clue Brier | Gemini (0.0552) | Exceptional calibration with citations |
| 16 | Worst With Clue Brier | DeepSeek (0.2978) | Dangerous overconfidence persists |
15. CONCLUSION
Silent Failure — confident wrongness without warning — represents a significant risk in legal AI deployment. Across 960 interactions, 222 Silent Failures occurred, representing 23.1% of all predictions. The risk varied by jurisdiction: Australia recorded the highest Silent Failure rate (30.8%), followed by Nigeria (23.8%), the United States (22.5%), and the United Kingdom (15.4%).
DeepSeek accounted for 59 Silent Failures (26.6% of the total), with its highest observed burdens in Australia (20), the United States (19), and Nigeria (17). Gemini showed the lowest observed risk, with only 23 Silent Failures and no Silent Failures observed in the United States.
Critically, citation context reduced Silent Failures by 34.3% overall (134 → 88). Nigeria showed the largest improvement (-53.8%), confirming that simple prompt design changes materially improve model reliability, particularly in low resource jurisdictions.
Brier Score confirms the EV findings: DeepSeek had the poorest calibration quality observed (Brier = 0.2978 With Clue) and the highest Silent Failure burden (59), while Gemini achieved the strongest With Clue reliability profile (Brier = 0.0552, Silent Failures = 23).
Extreme cases in Nigeria and the United States produced EV = -0.95 to -1.0, demonstrating that models can be perfectly confident while being completely wrong. This pattern is invisible under conventional accuracy based benchmarking.
For practitioners: always include citation context, exercise caution with DeepSeek outputs, and treat Australian law as the jurisdiction with the highest observed Silent Failure risk. The evidence is clear: Silent Failures are measurable, jurisdiction dependent, and reducible through prompt design. This is the empirical foundation for evidence based legal AI deployment.
16. RECOMMENDATIONS
| Recommendation | Evidence |
|---|---|
| Adopt Silent Failure as a core evaluation metric | 222 Silent Failures identified (23.1%) |
| Always include citation context | 34.3% Silent Failure reduction |
| Exercise caution with DeepSeek for Australian, Nigerian, and US legal practice | 59 Silent Failures; multiple EV = -1.0 cases; Brier = 0.2978 |
| Consider Gemini for low risk deployment | 23 Silent Failures; 0 in US; Brier = 0.0552 |
| Implement jurisdiction specific risk protocols | Silent Failure rates vary 15.4%–30.8% |
| Verify high confidence outputs independently | Extreme cases show perfect confidence on wrong answers |
| Include Brier Score alongside accuracy in reporting | Clear association between Brier Score and Silent Failures |
| Test Grok both with and without citations | Grok's Silent Failures slightly increase with clues (+1) |
Report End
QUICK REFERENCE CARD
| Metric | Value |
|---|---|
| Total Silent Failures | 222 (23.1%) |
| Highest Jurisdiction | Australia (74, 30.8%) |
| Second Highest | Nigeria (57, 23.8%) |
| Lowest Jurisdiction | United Kingdom (37, 15.4%) |
| Highest Model | DeepSeek (59) |
| Lowest Model | Gemini (23) |
| Worst Cell | DeepSeek in Australia (20) |
| Best Cell | Gemini in US (0) |
| No Clue SFs | 134 |
| With Clue SFs | 88 |
| Reduction | 46 (34.3%) |
| Best With Clue Brier | Gemini (0.0552) |
| Worst With Clue Brier | DeepSeek (0.2978) |
| Critical Cases | EV = -0.95 to -1.00 |
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Expected Calibration Error by model and prompting condition
Model No Clue With Clue
Claude 0.0375 0.0959
Grok 0.0571 0.0256
Gemini 0.1459 0.0188
ChatGPT 0.1611 0.0980
Perplexity 0.1793 0.1100
DeepSeek 0.2898 0.2620Thesis Table 4.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Brier score by model and prompting condition
Model No Clue With Clue Relative change
Gemini 0.184 0.055 −70.1%
ChatGPT 0.239 0.156 −34.7%
Perplexity 0.237 0.164 −30.8%
Claude 0.187 0.138 −26.2%
DeepSeek 0.327 0.298 −8.9%
Grok 0.133 0.145 +9.0%Thesis Table 4.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Silent Failure summary across the primary benchmark
Measure Observed value
Total Silent Failures 222 of 960 (23.1%)
Under No Clue 134 (of 134 incorrect)
Under With Clue 88 (of 90 incorrect)
Reduction in count 34.3%
Highest per-model count DeepSeek (59)
Second-highest per-model count Perplexity (41)
Third-highest per-model count ChatGPT (39)Thesis Table 4.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Silent Failure rate by jurisdiction, aggregated across models and conditions (240 predictions each)
Jurisdiction Resource classification Rate Approximate count
Australia Mid-resource 30.8% 74
Nigeria Low-resource 23.8% 57
United States High-resource 22.5% 54
United Kingdom High-resource 15.4% 37
Counts are derived from the reported rates over 240 predictions per jurisdiction and sum to the
222 total.Thesis Table 4.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Selected high-confidence incorrect predictions
Case Jurisdiction Stated confidence Expected Value
NG 004 Nigeria 95% −0.95
NG 009 Nigeria 95% −0.95
US 008 United States 100% −1.00
US 010 United States 100% −1.00Thesis Table 4.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Cases predicted incorrectly by all six models under No Clue
No Clue With Clue
Case Jurisdiction Legal issue represented
correct correct
NG 009 Nigeria Procedural limitation 0/6 2/6
NG 020 Nigeria Document hierarchy 0/6 5/6
AU 013 Australia Purposive statutory construction 0/6 3/6Thesis Table 4.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.