CrossLaw/Part I · Benchmark/Confidence and Silent Failure
07 / 26 · Part I · BenchmarkThesis §§4.4–4.6; Tables 4.5–4.10

222 of 960 predictions were confident errors.

Confidence and Silent Failure

Key takeaways

  • Silent Failure is the study’s label for an incorrect verdict at >50% stated confidence.
  • The count fell from 134 No Clue to 88 With Clue.
  • Australia had the highest jurisdiction rate, 30.8%.

What was tested & why

Confidence and Silent Failure

Accuracy and confidence reliability are separable. Expected Calibration Error and Brier score are shown in the thesis tables; high confidence is not an independent check of correctness.

Evidence boundary

Silent Failure is a study-specific operational measure, not a general industry classification.

Source: Thesis §§4.4–4.6; Tables 4.5–4.10 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Silent Failure by jurisdiction

Incorrect verdict at >50% confidence
Australia
30.8%
Nigeria
23.8%
United States
22.5%
United Kingdom
15.4%

Thesis Table 4.8 · 240 predictions per jurisdiction. Percentages are source-reported; approximate counts are in the table.

Model-level comparison

All six models, side by side

Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.

Historical report Table 3 · model × jurisdiction · report-only

Silent Failures by model and jurisdiction (both conditions, 40 predictions per cell)

ModelNigeriaAustraliaUSUKTotal /160
ChatGPT61413639
Gemini970723
Claude9117633
Grok7122627
DeepSeek172019359
Perplexity91013941
All models57745437222

Report-only cell counts. Cross-check: column totals equal thesis Table 4.8 (57, 74, 54, 37), the grand total equals Table 4.7 (222), and the DeepSeek, Perplexity and ChatGPT totals equal Table 4.7’s three highest model counts. The thesis does not publish the individual cells. DeepSeek’s burden was concentrated outside the UK; Gemini had none on US cases.

Historical report Table 5 · model × condition · report-only

Silent Failures by model and prompting condition (80 predictions per cell)

ModelNo ClueWith ClueChangeIncorrect NCIncorrect WC
ChatGPT2514-112514
Gemini185-13185
Claude2112-92112
Grok131411316
DeepSeek3227-53227
Perplexity2516-92516
All models13488-4613490

Silent Failure counts are report-only; condition totals (134 and 88, a 34.3% reduction) match thesis Table 4.7. Incorrect counts are 80 minus correct from Tables 4.1–4.2: under No Clue every error was a Silent Failure; under With Clue only two Grok errors were not. Grok was the only model with more Silent Failures after the clue.

Historical report Table 4 · jurisdiction × condition · report-only

Silent Failures by jurisdiction and condition (all six models)

GroupNo ClueWith ClueReduction% reduction
Nigeria39182153.8
Australia43311227.9
United States2925413.8
United Kingdom2314939.1
Total134884634.3

Report-only split; totals match thesis Tables 4.7–4.8. Shown beside the model tables so the jurisdiction view is not read alone.

Thesis Tables 4.5–4.6 · model calibration · source-reported

Calibration by model: Brier score and Expected Calibration Error (lower is better)

ModelBrier NCBrier WCBrier change %ECE NCECE WC
ChatGPT0.2390.156-34.70.16110.098
Gemini0.1840.055-70.10.14590.0188
Claude0.1870.138-26.20.03750.0959
Grok0.1330.14590.05710.0256
DeepSeek0.3270.298-8.90.28980.262
Perplexity0.2370.164-30.80.17930.11

Thesis values. DeepSeek had the weakest Brier and ECE in both conditions; Gemini had the lowest With Clue Brier and ECE; Grok had the lowest No Clue Brier and was the only model whose Brier score rose with the clue. Change column is not shaded.

Benchmarking report · report-only · full record

What the benchmarking report adds to this page

These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.

Section 13 · Notable failure casesReport lines 677–702

SECTION 13: NOTABLE FAILURE CASES

13.1 NG_009 – Ethiopian Airlines v Polaris Bank (Nigeria)

ConditionCorrectIssue
No Clue0/6All models reasoned on merits, missed statute-bar entirely
With Clue2/6Only ChatGPT and Gemini corrected

Lesson: LLMs systematically underperform on procedural/limitation points – a critical risk for practitioners. This demonstrates a structural bias toward substantive merits reasoning over procedural analysis. Even when the correct legal outcome depends entirely on a limitation period, models default to analyzing the substantive merits of the dispute.

13.2 NG_020 – Atiba v Suberu (Nigeria)

ConditionCorrectIssue
No Clue0/6All models predicted borrower's appeal would succeed, assuming a letter superseded a formal deed
With Clue5/6All except DeepSeek corrected with citation

Lesson: The parol evidence / document hierarchy problem is extremely common in Nigerian property and banking disputes. All six models defaulted to the wrong intuition – that a subsequent letter from the lender modified the formal mortgage deed. This tells practitioners something specific about where not to trust these tools without verification. The universal failure on this case is as significant as NG_009, revealing a systematic weakness in understanding document hierarchy and integration rules.

13.3 AU_013 – Ecosse Property Holdings (Australia)

ConditionCorrectIssue
No Clue0/6All models applied literal textual construction of "payable by the tenant"
With Clue3/6ChatGPT, Gemini, Grok corrected with citation

Lesson: Purposive commercial construction is a systematic weakness; citation retrieval helps but is not universal. This case represents the only universal failure in the Australian dataset and demonstrates that models default to literal interpretation even when commercial purpose demands a different construction.

13.4 UK_004 – G4S v Lewis-Ranwell (UK)

ConditionCorrectIssue
With Clue3/6Failed by ChatGPT, Grok, DeepSeek

Lesson: Illegality defence application to insanity verdicts is genuinely difficult – even with citation. This 2026 judgment tests training data frontiers and shows that even when citations are provided, models may struggle with novel or complex policy questions.

13.5 US_013 – Medical Marijuana v Horn (US)

ConditionCorrectIssue
No Clue3/6Only 3 models correct
With Clue2/6Only Gemini and Grok correct

Lesson: Most recent case (2025) tests training data frontier – significant recency effect. Even with citation, 4 of 6 models failed, demonstrating that models cannot be relied upon for very recent jurisprudence.

Report Appendix A · Supplementary statistical analyses (logistic regression, ECE, domain chi-square, power)Report lines 2455–2592

APPENDIX A: Supplementary Statistical Analyses Report (Optional)

1. Introduction and Rationale

The primary statistical validation (Part 2 of the main report) established key findings using non parametric tests (Friedman, McNemar, Mann Whitney U, Jonckheere Terpstra, Chi square, Cochran’s Q, coefficient of variation, and reliability diagrams). Those tests were appropriate for the experimental design and provided robust evidence for the headline conclusions.

However, to further strengthen the research and address potential peer review questions, three supplementary statistical evaluations were conducted:

  • Mixed effects logistic regression – to quantify the simultaneous effect of Condition, Model, and Legal Domain on accuracy while controlling for case level variance (a main effects logit model was used as an approximation).
  • Expected Calibration Error (ECE) – to provide a single numeric measure of confidence calibration for each model and condition, complementing the visual reliability diagrams.
  • Domain specific Chi square tests – to test whether accuracy varies significantly across legal domains overall and per jurisdiction.
  • Power analysis for non significant results – to determine if the sample size was adequate to detect meaningful effects for the non significant findings (e.g., Grok’s decline with clues, Claude’s marginal improvement).

These analyses are optional but add rigour, confirm the adequacy of the sample size, quantify calibration, and demonstrate the influence of legal domain. They are placed in this Supplementary Statistical Analyses section, which can be appended to Part 2 of the main report as Appendix A.

2. Methodology

All analyses were performed using the complete dataset of 960 verified experiments (20 cases × 4 jurisdictions × 6 models × 2 conditions). Data preparation followed the same protocol as the main study (sanitised facts, binary verdict scores, confidence scores scaled to [0,1]). The code was executed in a Google Colab environment using Python with libraries statsmodels, scikit-learn, pingouin, and numpy.

2.1 Logistic Regression (Main Effects)

A logistic regression model was fitted with the binary outcome correct (1 = correct verdict, 0 = incorrect). Fixed effects included:

  • Condition_num (0 = No Clue, 1 = With Clue)
  • Model (categorical, reference = ChatGPT)
  • Legal_Domain (categorical, reference = not explicitly set; the output shows comparisons for many domain levels)

The model was estimated using maximum likelihood (method='bfgs') with 1000 iterations to ensure convergence. No random effects were included due to the complexity of crossed random effects in Python; the model serves as an approximation of the main fixed effects. Consequently, within case correlation is not accounted for, but the fixed effect coefficients are consistent with the non parametric tests.

2.2 Expected Calibration Error (ECE)

ECE quantifies the average absolute difference between predicted confidence and observed accuracy across bins. A lower ECE indicates better calibration.

For each model–condition pair, predicted confidences were divided into 10 equal width bins [0,0.1), [0.1,0.2), … , [0.9,1.0]. For each bin, the mean confidence and observed accuracy were calculated. ECE is the weighted average of the absolute differences:

"ECE"=∑_(b=1)^B▒n_b/N∣〖"acc" 〗_b-〖"conf" 〗_b∣

where n_b is the number of predictions in bin b, N total predictions, 〖"acc" 〗_b observed accuracy, and 〖"conf" 〗_b mean confidence. Empty bins were safely ignored.

2.3 Domain specific Chi square Tests

For each jurisdiction and for all jurisdictions combined, a contingency table of Legal_Domain vs. correct was constructed. A chi square test of independence was performed to determine whether the distribution of correct/incorrect answers differs significantly by legal domain. A significant result indicates that some legal domains are systematically easier or harder for LLMs.

2.4 Power Analysis for Non significant Results

Two non significant results from the main McNemar tests were examined:

  • Grok’s decline with clues (−2.5%, p = 0.8238)
  • Claude’s improvement (+11.25%, p = 0.0784)

For each, the effect size Cohen’s h was calculated:

h=2arcsin⁡(√(p_1 ))-2arcsin⁡(√(p_2 ))

where p_1 and p_2 are the proportions correct in the two conditions (No Clue vs With Clue). Using proportion_effectsize and tt_ind_solve_power (two sample independent proportions power) with n=80 per group, α=0.05, two tailed, the statistical power to detect the observed effect was estimated. This provides a check on whether non significance could be due to insufficient sample size.

Note: The independent samples approximation is conservative; the exact McNemar power depends on discordant pairs, but the directional conclusions remain valid.

3. Results

3.1 Logistic Regression (Main Effects)

The logistic regression converged successfully (Pseudo R squared = 0.2089, log likelihood ratio test p < 0.0001). The full coefficient table is presented below.

Table 1: Logistic Regression Results (Main Effects Only)

| Variable | Coefficient | Std. Error | z | P>|z| | 95% CI |

|----------|------------|------------|----|-------|--------|

| Intercept | 2.0927 | 1.081 | 1.935 | 0.053 | [-0.027, 4.212] |

| Condition_num | 0.6729 | 0.176 | 3.827 | 0.000 | [0.328, 1.018] |

| Model (ref: ChatGPT) | | | | | |

| Claude | 0.2678 | 0.299 | 0.895 | 0.371 | [-0.319, 0.854] |

| DeepSeek | -0.7591 | 0.279 | -2.719 | 0.007 | [-1.306, -0.212] |

| Gemini | 0.7973 | 0.322 | 2.477 | 0.013 | [0.167, 1.428] |

| Grok | 0.5159 | 0.309 | 1.672 | 0.095 | [-0.089, 1.121] |

| Perplexity | -0.0835 | 0.289 | -0.289 | 0.773 | [-0.650, 0.483] |

| Legal_Domain (selected significant) | | | | | |

| Commercial Law (Banking) | -4.2719 | 1.327 | -3.219 | 0.001 | [-6.873, -1.671] |

| Commercial Law (Ethics) | -3.3053 | 1.239 | -2.667 | 0.008 | [-5.734, -0.876] |

| Contract Law (Construction) | -3.3053 | 1.239 | -2.667 | 0.008 | [-5.734, -0.876] |

| Contract Law (Mortgage) | -2.9211 | 1.226 | -2.383 | 0.017 | [-5.324, -0.518] |

| Equity (Trusts) | -2.5565 | 1.222 | -2.092 | 0.036 | [-4.951, -0.161] |

| Civil Law (RICO) | -2.5565 | 1.222 | -2.092 | 0.036 | [-4.951, -0.161] |

| Intellectual Property (Trademark) | -2.0661 | 1.121 | -1.842 | 0.065 | [-4.264, 0.132] |

Note: Only a subset of legal domains with p < 0.10 are shown for brevity; the full table is available in the code output.

Key findings:

  • The Condition_num coefficient (+0.6729, p < 0.001) confirms that including citations significantly increases the log odds of a correct verdict, consistent with the McNemar test.
  • DeepSeek is significantly worse than ChatGPT (p = 0.007), while Gemini is significantly better (p = 0.013). Grok is marginally better (p = 0.095).
  • Several legal domains (e.g., Commercial Law (Banking), Contract Law (Construction), Equity (Trusts)) show large negative coefficients, indicating that models perform worse on those domains relative to the reference category.

3.2 Expected Calibration Error (ECE)

Table 2 reports the ECE for each model and condition. Lower ECE indicates better alignment between confidence and observed accuracy.

ECE < 0.05 is excellent, 0.05–0.10 is good, > 0.20 indicates serious miscalibration.

Table 2: Expected Calibration Error by Model and Condition

ModelConditionECE
ClaudeNo Clue0.0375
GrokNo Clue0.0571
GeminiNo Clue0.1459
ChatGPTNo Clue0.1611
PerplexityNo Clue0.1793
DeepSeekNo Clue0.2898
GeminiWith Clue0.0188
GrokWith Clue0.0256
ClaudeWith Clue0.0959
ChatGPTWith Clue0.0980
PerplexityWith Clue0.1100
DeepSeekWith Clue0.2620

Interpretation:

  • Best calibrated: Claude (No Clue, ECE=0.0375) and Gemini (With Clue, ECE=0.0188) – their confidences reflect actual accuracy almost perfectly.
  • Worst calibrated: DeepSeek in both conditions (ECE = 0.290 and 0.262) – it is dangerously overconfident, frequently expressing high confidence on incorrect predictions. This confirms the “Silent Failure” risk identified in the descriptive analysis.
  • Grok has low ECE in both conditions, indicating good calibration despite its accuracy decline with clues.
  • ChatGPT and Perplexity show moderate miscalibration, improving with clues.

3.3 Domain specific Chi square Tests

Table 3: Chi square Test for Independence of Legal Domain and Correctness

Jurisdictionχ²dfp-valueSignificance (α=0.05)
Overall (all jurisdictions)162.994730.0000Significant
Nigeria65.388150.0000Significant
United States25.942140.0173Significant
Australia45.992160.0003Significant
United Kingdom22.080180.2803Not significant

Findings:

  • Accuracy varies significantly by legal domain overall and in Nigeria, the United States, and Australia.
  • In the United Kingdom, domain was not a significant factor (p = 0.2803). This suggests that, uniquely among the four jurisdictions, UK legal domains are handled more uniformly by the models. This may reflect the high quality centralised reporting of UK case law (BAILII) and/or DeepSeek’s memorisation of UK cases, which reduces domain specific variation.

3.4 Power Analysis for Non significant Results

Table 4: Power to Detect Observed Effect Sizes (Independent proportions approximation)

ModelComparisonp value (McNemar)p₁ (NC)p₂ (WC)Cohen’s hEffect sizePower (two sample independent)
GrokNC vs WC0.82380.83750.80000.097Negligible0.094
ClaudeNC vs WC0.07840.73750.8500–0.280Small medium0.422

Interpretation:

  • Grok: The observed decline (Cohen’s h = 0.097) is negligible. The power to detect such a small effect with 80 cases is only 9.4%, meaning the non significant result is not due to low power but because the true effect is practically zero. The descriptive finding that Grok declines with clues is not statistically meaningful.
  • Claude: The improvement (Cohen’s h = 0.280) is a small to medium effect. The power of 42% indicates that the current sample size had only a moderate chance of detecting this effect. Hence the marginal p value (0.0784) could be a false negative – the improvement may be real but requires a larger sample to confirm. This supports the descriptive observation that Claude’s improvement (+11.25%) is consistent across jurisdictions and may be genuine.
  • Overall sample size (80 cases per condition) is adequate to detect a medium effect (Cohen’s h ≥ 0.3) with power > 0.8 (as shown in the output). Non significant results with smaller effects are therefore not due to insufficient data.

Note: These power calculations use an independent samples proportion test, a conservative approximation for the paired McNemar design. Actual power for McNemar would be slightly higher, but the directional conclusions remain valid.

4. Integration with Prior Findings

4.1 Consistency with Descriptive and Main Statistical Analysis

  • Condition effect (+9.1% average improvement) is strongly confirmed by the logistic regression (p < 0.001) and the low ECE of Gemini with clues.
  • DeepSeek’s poor performance is reinforced: logistic regression shows it is significantly worse than ChatGPT (p = 0.007), and its high ECE (>0.26) confirms overconfidence, aligning with the “Silent Failure” warning in the descriptive section.
  • Gemini’s excellence with clues is supported by the lowest ECE (0.0188) and the significant positive coefficient (p = 0.013).
  • Grok’s clue induced decline is shown to be negligible (h = 0.097) and not statistically meaningful; however, its calibration remains good (ECE = 0.0256 with clues). This clarifies that while Grok’s accuracy does not improve, it does not become dangerously miscalibrated.
  • Domain variability was already observed descriptively (e.g., NG_020 and AU_013 domain failures). The chi square tests quantify this: accuracy differs by domain in three of four jurisdictions, and the UK is the exception.

4.2 Contributions of Supplementary Analyses

Supplementary TestAdded Value
Logistic regressionQuantifies the simultaneous effect of Condition, Model, and Domain, confirming the strength of citation context and identifying domains of systematic difficulty.
Expected Calibration ErrorProvides a single numeric calibration metric, objectively showing that DeepSeek is dangerously overconfident, while Gemini and Claude are well calibrated.
Domain chi squareEmpirically demonstrates that legal domain significantly impacts performance, with the UK being uniquely uniform – a new finding that may reflect data quality.
Power analysisReassures that non significant results (Grok’s decline) are truly negligible, while Claude’s marginal improvement may warrant further study.

4.3 Implications for the Industrial Translation

  • ECE reinforces the “Silent Failure” warning for DeepSeek: not only is its accuracy low, but its confidence is systematically misleading. The practitioner’s Duty of Inquiry Checklist must mandate independent verification of DeepSeek outputs regardless of confidence.
  • Domain chi square identifies specific legal domains (e.g., Commercial Law (Banking), Contract Law (Construction), Equity (Trusts)) as high risk areas where extra caution is required.
  • Power analysis provides evidence that the sample size (80 cases × 2 conditions) is sufficient for medium sized effects, strengthening the credibility of the null findings (e.g., Grok’s decline).

5. Placement in the Overall Report

These supplementary analyses are not required for the main conclusions but provide valuable depth and address potential methodological questions from reviewers. They should be placed as:

6. Conclusion

The supplementary statistical analyses confirm and extend the primary findings:

  • Citation context significantly improves accuracy (logistic regression p < 0.001).
  • DeepSeek is both low accuracy and severely overconfident (ECE > 0.26), validating the “Silent Failure” risk.
  • Gemini and Claude are well calibrated, with Gemini achieving near perfect calibration with clues (ECE = 0.0188).
  • Legal domain affects performance in Nigeria, US, and Australia, but not in the UK – a novel observation that may reflect training data characteristics.
  • The sample size (80 cases) is sufficient to detect medium effects; Grok’s non significant decline is truly negligible, while Claude’s marginal improvement may merit further investigation.

These optional analyses add rigour, address potential critique, and provide deeper insight into model behaviour. They are now ready for inclusion in the final thesis or paper.

SILENT FAILURE ANALYSIS AND FINDINGS — CROSSLAW BENCHMARK

A Comprehensive Analysis of Confidently Wrong LLM Predictions Across Four Common Law Jurisdictions

Silent Failure analysis and findings reportReport lines 2593–2907

Prepared for: HUMN4002 RD2 Viva Voce | Western Sydney University

Date: Autumn 2026

Analytical Framework: Silent Failure = EV < -0.5 (Incorrect Verdict + High Confidence)

Data Source: 960 verified interactions (80 cases × 6 models × 2 conditions)

EXECUTIVE SUMMARY

Silent Failure — confident wrongness without warning — represents a significant risk in legal AI deployment. A model that is wrong but expresses high confidence creates outputs that appear authoritative and persuasive to practitioners, making errors difficult to detect and correct.

Across 960 controlled interactions, we identified 222 Silent Failures, representing 23.1% of all predictions. Citation context reduced Silent Failures by 34.3% (134 → 88), confirming that prompt design materially improves model reliability.

Australia recorded the highest Silent Failure burden within the benchmark (74, 30.8%), followed by Nigeria (57, 23.8%), United States (54, 22.5%), and United Kingdom (37, 15.4%). DeepSeek accounted for the largest model-specific burden (59 Silent Failures), with its highest observed burdens in Australia (20), the United States (19), and Nigeria (17). Gemini demonstrated the strongest Silent Failure safety profile among the models tested, producing only 23 Silent Failures total and zero in the United States.

Brier Score confirms the EV findings: DeepSeek had the poorest calibration quality observed (Brier = 0.2978 With Clue) and the highest Silent Failure burden (59), while Gemini achieved the strongest With Clue reliability profile (Brier = 0.0552, Silent Failures = 23).

Critical cases in Nigeria and the United States produced EV = -0.95 to -1.0, demonstrating that models can be perfectly confident while being completely wrong — a pattern that would remain invisible under conventional accuracy based benchmarking.

1. DEFINITION AND METHODOLOGY

1.1 Silent Failure Definition

Accuracy alone does not capture the practical risks associated with deploying LLMs in legal contexts. A model may produce an incorrect answer while expressing very high confidence, creating a particularly dangerous form of error that can appear persuasive to practitioners.

Following the benchmark framework, a Silent Failure was defined as any prediction with an Expected Value (EV) below −0.5.

ConditionThreshold
Verdict Score0 (incorrect)
Expected Value (EV)< -0.5 (confidently wrong)

1.2 Expected Value Formula

The Expected Value metric combines accuracy and confidence into a single risk adjusted measure:

OutcomeFormula
Correct PredictionEV = +Confidence / 100
Incorrect PredictionEV = −Confidence / 100

Interpretation:

EV RangeMeaning
EV > 0.5Reliable prediction (correct with moderate confidence)
0 < EV < 0.5Correct but uncertain
-0.5 < EV < 0Incorrect but appropriately cautious
EV < -0.5⚠️ SILENT FAILURE — Confidently wrong; worse than random guessing

1.3 Brier Score Definition

Brier Score measures the mean squared difference between predicted probabilities (Confidence/100) and actual outcomes (0 or 1). Lower Brier Scores indicate better calibration.

Note on Brier Score Interpretation: The thresholds used in this report (e.g., 0.0–0.1 as "excellent") are heuristic guidelines for comparative analysis within this benchmark rather than universally accepted calibration thresholds for legal AI systems. They are intended to facilitate model comparison rather than establish absolute standards.

Brier Score RangeHeuristic Interpretation
0.0 – 0.1Excellent (relative to benchmark)
0.1 – 0.2Good (relative to benchmark)
0.2 – 0.3Moderate (relative to benchmark)
> 0.3Poor (relative to benchmark)

1.4 Dataset Overview

ParameterValue
Total Cases80
JurisdictionsNigeria, Australia, United Kingdom, United States
ModelsChatGPT 4.1, Gemini 1.5 Pro, Claude 3 Opus, Grok 2, DeepSeek V3, Perplexity Sonar
Prompt ConditionsNo Clue (facts only), With Clue (facts + citation + jurisdiction + year + domain)
Total Interactions960 (80 total cases × 6 models × 2 conditions)
Silent Failure DefinitionEV < -0.5

2. OVERALL SILENT FAILURE BURDEN

Across the 960 benchmark interactions, 222 predictions met the Silent Failure threshold, representing 23.1% of all model outputs.

This finding indicates that almost one in four legal AI predictions within the benchmark were confidently incorrect enough to satisfy the Silent Failure criterion. The result demonstrates that accuracy alone substantially understates deployment risk.

3. SILENT FAILURE BURDEN ACROSS JURISDICTIONS

Table 1. Silent Failures by Jurisdiction

JurisdictionSilent FailuresTotal PredictionsRate
Australia7424030.8%
Nigeria5724023.8%
United States5424022.5%
United Kingdom3724015.4%
TOTAL22296023.1%

Key Observations

  • Australia recorded the highest Silent Failure rate within the benchmark (30.8%) — almost one in three predictions were confidently wrong. This suggests that the evaluated models were less reliable on Australian appellate law tasks than on the other jurisdictions examined.
  • Nigeria recorded the second highest rate (23.8%) — nearly one in four predictions were Silent Failures. This confirms that low resource jurisdictions face compounding risks.
  • The United Kingdom recorded the lowest Silent Failure rate (15.4%) — reflecting better training data coverage and model calibration for UK law.
  • The UK was the only jurisdiction below 20% — suggesting that high resource, well digitised legal systems benefit from more reliable AI calibration.

4. SILENT FAILURE DISTRIBUTION BY MODEL

Table 2. Silent Failure Counts by Model

ModelSilent Failures% of Total
DeepSeek5926.6%
Perplexity4118.5%
ChatGPT3917.6%
Claude3314.9%
Grok2712.2%
Gemini2310.4%
TOTAL222100%

Key Observations

  • DeepSeek accounted for 26.6% of all Silent Failures (59 of 222) — DeepSeek was the highest risk model in terms of calibration failure within this benchmark.
  • Gemini had the lowest observed Silent Failure burden (23) — Gemini was the lowest risk model among those tested.
  • Grok had 27 Silent Failures despite high accuracy (83.8%) — Grok was accurate when correct but showed overconfidence when wrong.
  • Perplexity's high Silent Failure count (41) — retrieval first models may be more prone to overconfidence when retrieval is incorrect.

5. SILENT FAILURE HEATMAP (MODEL × JURISDICTION)

Table 3. Silent Failure Counts by Model and Jurisdiction

ModelNigeriaAustraliaUnited StatesUnited KingdomTotal
ChatGPT61413639
Gemini970723
Claude9117633
Grok7122627
DeepSeek172019359
Perplexity91013941
Total57745437222

Key Observations

FindingInsight
DeepSeek in Australia had the highest observed cell burden (20)Australian practitioners face elevated risk when using DeepSeek
DeepSeek in the United States was second highest (19)US practitioners also face elevated Silent Failure risk with DeepSeek
DeepSeek in Nigeria was also high (17)Nigerian practitioners face serious risk with DeepSeek deployment
Gemini in the US had zero Silent Failures (0)Gemini was exceptionally well calibrated for US law within this benchmark
Gemini had the lowest total Silent Failures (23)Gemini was the lowest risk model overall
DeepSeek in the UK had the lowest cell burden (3)DeepSeek's UK specialised knowledge was well calibrated, but this does not indicate safety elsewhere

6. IMPACT OF CITATION CONTEXT ON SILENT FAILURE RISK

One of the most important findings concerns the effect of citation context.

Table 4. Silent Failures by Condition and Jurisdiction

JurisdictionNo ClueWith ClueReduction% Reduction
Nigeria39182153.8%
Australia43311227.9%
United States2925413.8%
United Kingdom2314939.1%
TOTAL134884634.3%

Key Observations

  • Citation context reduced Silent Failures by 34.3% overall (134 → 88). This is one of the strongest findings: prompt design materially improved model reliability.
  • Nigeria showed the largest absolute reduction (39 → 18, -53.8%) — citation clues most helped Nigerian legal reasoning, where models otherwise struggled.
  • The United Kingdom showed strong reduction (23 → 14, -39.1%) — citation context significantly improved UK law calibration.
  • The United States showed the smallest reduction (29 → 25, -13.8%) — clues did not substantially reduce Silent Failures in US law, possibly because models were already overconfident in US legal contexts.
  • Australia showed moderate improvement (43 → 31, -27.9%) — clues helped but did not fully resolve Australian calibration issues.

7. SILENT FAILURE BY MODEL × CONDITION

Table 5. Model Specific Silent Failures by Condition

ModelNo ClueWith ClueDifference% Reduction
ChatGPT2514-1144.0%
Gemini185-1372.2%
Claude2112-942.9%
Grok1314+1-7.7% (worsens)
DeepSeek3227-515.6%
Perplexity2516-936.0%
TOTAL13488-4634.3%

Key Observations

  • Every model except Grok improved with citation context — Grok showed a slight increase in Silent Failures with clues (+1), consistent with its accuracy decline.
  • Gemini showed the largest improvement (-13, 72.2%) — citation context dramatically improved Gemini's calibration.
  • DeepSeek showed the smallest improvement (-5, 15.6%) — DeepSeek remained dangerously overconfident even with clues.
  • Gemini had the lowest Silent Failure count in both conditions (18 No Clue, 5 With Clue) — lowest risk model overall.

8. BRIER SCORE CALIBRATION FINDINGS

Brier Score measures the mean squared difference between predicted probabilities (Confidence/100) and actual outcomes (0 or 1). Lower Brier Scores indicate better calibration.

Table 6. Average Brier Score by Model and Condition

ModelNo Clue BrierWith Clue BrierInterpretation
ChatGPT0.23870.1559Improves with clues
Gemini0.18400.0552Best With Clue
Claude0.18700.1375Improves with clues
Grok0.13280.1447Best No Clue; slight decline with clues
DeepSeek0.32740.2978Worst in both conditions
Perplexity0.23680.1642Improves with clues

Key Observations

  • Gemini achieved the best With Clue calibration observed (Brier = 0.0552) — exceptionally well calibrated when citation context was provided.
  • Grok achieved the best No Clue calibration observed (Brier = 0.1328) — well calibrated without clues, consistent with its stable performance.
  • DeepSeek had the poorest calibration observed in both conditions — Brier > 0.29 in both No Clue and With Clue, confirming dangerous overconfidence.
  • All models except Grok showed improved calibration with citation context — Grok's Brier slightly worsened (0.1328 → 0.1447), consistent with its accuracy decline.

9. RELATIONSHIP BETWEEN BRIER SCORE, EXPECTED VALUE, AND SILENT FAILURE RISK

Although Brier Score and Silent Failure measure different aspects of model behaviour, the two metrics were broadly aligned in this benchmark. A clear association was observed between poorer calibration (higher Brier Scores) and higher Silent Failure counts. Models with poorer calibration generally produced more Silent Failures, suggesting that calibration quality appears to be a useful indicator of deployment risk within this benchmark.

Table 7. Brier Score vs Silent Failures by Model

ModelWith Clue BrierSilent FailuresInterpretation
DeepSeek0.297859Poorest calibration + highest Silent Failures
Perplexity0.164241Moderate calibration
ChatGPT0.155939Good calibration
Claude0.137533Good calibration
Grok0.144727Moderate calibration (slightly worsens with clues)
Gemini0.055223Best calibration + lowest Silent Failures

Key Observations

  • A clear association is observed between poorer calibration (higher Brier Scores) and higher Silent Failure counts.
  • DeepSeek showed the poorest calibration observed (Brier = 0.2978) and the highest Silent Failure burden (59).
  • Gemini showed the best calibration observed (Brier = 0.0552) and the lowest Silent Failure burden (23).
  • Grok showed good No Clue calibration but slightly worsened with clues, consistent with its unique decline pattern.
  • Brier Score confirms the EV findings — the same models that were overconfident (high Brier) were also the highest risk in terms of Silent Failure burden.

10. CRITICAL SILENT FAILURE CASES

Several cases demonstrated extreme Silent Failure behaviour.

Table 8. Example Silent Failure Cases (EV < -0.8)

CaseCase TitleJurisdictionModelConditionVerdictConfidenceEV
NG_004F.H.A. v. OyedejiNigeriaDeepSeekNo Clue095%-0.95
NG_009Ethiopian Airlines v. Polaris BankNigeriaDeepSeekNo Clue095%-0.95
NG_020Atiba Iyalamu v. SuberuNigeriaDeepSeekNo Clue095%-0.95
US_008Amgen Inc. v. SanofiUnited StatesDeepSeekNo Clue0100%-1.00
US_010Bartenwerfer v. BuckleyUnited StatesDeepSeekNo Clue0100%-1.00
UK_004G4S Health Services v. Lewis RanwellUnited KingdomDeepSeekWith Clue0100%-1.00

In each case, the model produced an incorrect legal prediction while maintaining near maximal or perfect confidence. These outputs would likely appear persuasive to practitioners despite being wrong.

Key Observations:

  • DeepSeek in Nigeria produced EV = -0.95 on multiple cases — near maximal confidence on incorrect predictions.
  • DeepSeek in the United States produced EV = -1.00 on two cases — 100% confidence on wrong answers.
  • DeepSeek in the United Kingdom produced EV = -1.00 on one With_Clue case — citation context did not prevent the Silent Failure.

Such cases illustrate why Silent Failure analysis provides information that traditional accuracy metrics cannot capture.

11. VISUALISATION READY SUMMARY

Silent Failure Risk by Jurisdiction

text

Australia ████████████████████████████████████████████ 30.8% (74/240)

Nigeria ██████████████████████████████████████████ 23.8% (57/240)

United States ████████████████████████████████████████ 22.5% (54/240)

United Kingdom █████████████████████████████████████████ 15.4% (37/240)

Silent Failure Burden by Model

text

DeepSeek ████████████████████████████████████████████ 59

Perplexity ████████████████████████████████████████████ 41

ChatGPT ████████████████████████████████████████████ 39

Claude ████████████████████████████████████████████ 33

Grok ████████████████████████████████████████████ 27

Gemini ████████████████████████████████████████████ 23

Condition Effect

text

No Clue ████████████████████████████████████████████████████████████████████████████████████ 134

With Clue ████████████████████████████████████████████████████████████████████████████████████ 88

↓ 34.3% reduction (46 fewer Silent Failures)

Brier Score by Model (With Clue)

text

DeepSeek ████████████████████████████████████████████ 0.2978

Grok ████████████████████████████████████████████ 0.1447

Perplexity ████████████████████████████████████████████ 0.1642

ChatGPT ████████████████████████████████████████████ 0.1559

Claude ████████████████████████████████████████████ 0.1375

Gemini ████████████████████████████████████████████ 0.0552

12. PRACTITIONER IMPLICATIONS

12.1 Model Selection Guidance

Use CaseRecommended ModelRationale
Lowest observed Silent Failure RiskGeminiOnly 23 Silent Failures total; 0 in US; Brier = 0.0552
Highest observed Silent Failure RiskDeepSeek59 Silent Failures; poorest calibration (Brier = 0.2978)
Australian LawExercise caution with DeepSeek20 Silent Failures in Australia — highest cell burden
Nigerian LawExercise caution with DeepSeek17 Silent Failures in Nigeria — critically high
US LawGeminiNo Silent Failures observed in the US benchmark
UK LawGemini or ClaudeLow Silent Failure rates; Gemini (7), Claude (6)

12.2 Prompt Design Guidance

InsightImplication
Citation context reduces Silent Failures by 34.3%Always include case citation, jurisdiction, year, and legal domain
Nigeria shows largest improvement (-53.8%)Nigerian practitioners benefit most from citation anchoring
US shows smallest improvement (-13.8%)US practitioners should not rely solely on citations to fix calibration
DeepSeek improves only 15.6% with cluesDeepSeek remains high risk even with optimal prompting

12.3 Brier Score and Calibration Guidance

InsightImplication
Gemini Brier = 0.0552 (With Clue)Gemini's confidence levels are well calibrated when citations are provided
DeepSeek Brier = 0.2978 (With Clue)DeepSeek's confidence is unreliable; always verify independently
Grok Brier = 0.1328 (No Clue)Grok is well calibrated without clues; confidence is more trustworthy
Brier Score association with Silent FailuresUse Brier Score as a proxy for deployment safety

12.4 DeepSeek Risk Warning

⚠️ Exercise caution when deploying DeepSeek in Australian, Nigerian, and United States legal contexts and require independent verification of outputs.

Evidence:

MetricValue
Total Silent Failures59 — highest of any model
Silent Failures in Australia20 — highest cell burden
Silent Failures in the United States19 — second highest cell burden
Silent Failures in Nigeria17 — critically high
With Clue ImprovementOnly 15.6% — calibration remained dangerously poor
With Clue Brier Score0.2978 — poorest calibration observed
Extreme CasesUS_008 (EV = -1.00), US_010 (EV = -1.00), UK_004 (EV = -1.00)

13. THEORETICAL CONTRIBUTIONS

13.1 Silent Failure Quantification

This study provides a systematic quantification of Silent Failures across multiple common law jurisdictions. Prior work identified Silent Failures conceptually (Pathak et al., 2025; Potts & Sudhof, 2026), but did not measure them across jurisdictions, models, and prompt conditions in a controlled benchmark.

13.2 Jurisdictional Risk Mapping

The finding that Australia recorded the highest Silent Failure rate (30.8%) while the UK recorded the lowest (15.4%) confirms that Silent Failure risk is jurisdiction dependent within this benchmark. This challenges the assumption that Silent Failure is a uniform model property.

13.3 Prompt Design as Risk Mitigation

The finding that citation context reduces Silent Failures by 34.3% confirms that prompt design materially improves calibration. This is one of the strongest practical findings: simple changes to input structure reduce the most dangerous failure mode.

13.4 Model Specific Risk Profiles

The identification of DeepSeek as the highest Silent Failure model and Gemini as the lowest provides a model specific Silent Failure risk taxonomy for legal AI within this benchmark framework.

13.5 Brier Score as an Indicator of Silent Failure Risk

The clear association observed between Brier Score and Silent Failure burden suggests that calibration quality may be a useful indicator of deployment safety. Models with poorer calibration (high Brier) consistently produced more Silent Failures within this benchmark.

13.6 Extreme Cases as Evidence of Overconfidence Risk

The existence of predictions with EV = -1.0 (100% confidence on incorrect verdicts) is consistent with concerns raised in the literature regarding overconfident language model outputs (Bender et al., 2021). These cases demonstrate that perfect confidence does not imply correct reasoning.

14. SUMMARY OF KEY FINDINGS

#FindingValueImplication
1Total Silent Failures222 (23.1%)Silent Failure is pervasive within the benchmark
2Highest Silent Failure JurisdictionAustralia (74, 30.8%)Australian appellate law tasks showed highest risk
3Second HighestNigeria (57, 23.8%)Low resource jurisdictions are vulnerable
4LowestUnited Kingdom (37, 15.4%)High resource systems showed lower risk
5Highest Silent Failure ModelDeepSeek (59)DeepSeek showed highest observed risk
6Lowest Silent Failure ModelGemini (23)Gemini showed lowest observed risk
7Worst CellDeepSeek in Australia (20)Highest risk combination identified
8Second Worst CellDeepSeek in US (19)US practitioners also face elevated risk
9Best CellGemini in US (0)No Silent Failures observed in the US benchmark
10No Clue Silent Failures134Overconfidence is worse without context
11With Clue Silent Failures88Citation context reduces risk
12Silent Failure Reduction46 (34.3%)Prompt design materially improves safety
13DeepSeek Extreme CasesEV = -1.0 (US_008, US_010)Perfect confidence on wrong answers
14Nigerian Critical CasesEV = -0.95 (NG_004, NG_009, NG_020)Near maximal confidence on wrong answers
15Best With Clue BrierGemini (0.0552)Exceptional calibration with citations
16Worst With Clue BrierDeepSeek (0.2978)Dangerous overconfidence persists

15. CONCLUSION

Silent Failure — confident wrongness without warning — represents a significant risk in legal AI deployment. Across 960 interactions, 222 Silent Failures occurred, representing 23.1% of all predictions. The risk varied by jurisdiction: Australia recorded the highest Silent Failure rate (30.8%), followed by Nigeria (23.8%), the United States (22.5%), and the United Kingdom (15.4%).

DeepSeek accounted for 59 Silent Failures (26.6% of the total), with its highest observed burdens in Australia (20), the United States (19), and Nigeria (17). Gemini showed the lowest observed risk, with only 23 Silent Failures and no Silent Failures observed in the United States.

Critically, citation context reduced Silent Failures by 34.3% overall (134 → 88). Nigeria showed the largest improvement (-53.8%), confirming that simple prompt design changes materially improve model reliability, particularly in low resource jurisdictions.

Brier Score confirms the EV findings: DeepSeek had the poorest calibration quality observed (Brier = 0.2978 With Clue) and the highest Silent Failure burden (59), while Gemini achieved the strongest With Clue reliability profile (Brier = 0.0552, Silent Failures = 23).

Extreme cases in Nigeria and the United States produced EV = -0.95 to -1.0, demonstrating that models can be perfectly confident while being completely wrong. This pattern is invisible under conventional accuracy based benchmarking.

For practitioners: always include citation context, exercise caution with DeepSeek outputs, and treat Australian law as the jurisdiction with the highest observed Silent Failure risk. The evidence is clear: Silent Failures are measurable, jurisdiction dependent, and reducible through prompt design. This is the empirical foundation for evidence based legal AI deployment.

16. RECOMMENDATIONS

RecommendationEvidence
Adopt Silent Failure as a core evaluation metric222 Silent Failures identified (23.1%)
Always include citation context34.3% Silent Failure reduction
Exercise caution with DeepSeek for Australian, Nigerian, and US legal practice59 Silent Failures; multiple EV = -1.0 cases; Brier = 0.2978
Consider Gemini for low risk deployment23 Silent Failures; 0 in US; Brier = 0.0552
Implement jurisdiction specific risk protocolsSilent Failure rates vary 15.4%–30.8%
Verify high confidence outputs independentlyExtreme cases show perfect confidence on wrong answers
Include Brier Score alongside accuracy in reportingClear association between Brier Score and Silent Failures
Test Grok both with and without citationsGrok's Silent Failures slightly increase with clues (+1)

Report End

QUICK REFERENCE CARD

MetricValue
Total Silent Failures222 (23.1%)
Highest JurisdictionAustralia (74, 30.8%)
Second HighestNigeria (57, 23.8%)
Lowest JurisdictionUnited Kingdom (37, 15.4%)
Highest ModelDeepSeek (59)
Lowest ModelGemini (23)
Worst CellDeepSeek in Australia (20)
Best CellGemini in US (0)
No Clue SFs134
With Clue SFs88
Reduction46 (34.3%)
Best With Clue BrierGemini (0.0552)
Worst With Clue BrierDeepSeek (0.2978)
Critical CasesEV = -0.95 to -1.00

Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 4.5 · source-reported

Expected Calibration Error by model and prompting condition

                           Model        No Clue     With Clue

                           Claude        0.0375       0.0959
                           Grok          0.0571       0.0256
                           Gemini        0.1459       0.0188
                           ChatGPT       0.1611       0.0980
                           Perplexity    0.1793       0.1100
                           DeepSeek      0.2898       0.2620

Thesis Table 4.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.6 · source-reported

Brier score by model and prompting condition

              Model          No Clue     With Clue     Relative change

              Gemini           0.184        0.055              −70.1%
              ChatGPT          0.239        0.156              −34.7%
              Perplexity       0.237        0.164              −30.8%
              Claude           0.187        0.138              −26.2%
              DeepSeek         0.327        0.298              −8.9%
              Grok             0.133        0.145              +9.0%

Thesis Table 4.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.7 · source-reported

Silent Failure summary across the primary benchmark

                Measure                               Observed value

                Total Silent Failures                 222 of 960 (23.1%)
                Under No Clue                         134 (of 134 incorrect)
                Under With Clue                       88 (of 90 incorrect)
                Reduction in count                    34.3%
                Highest per-model count               DeepSeek (59)
                Second-highest per-model count        Perplexity (41)
                Third-highest per-model count         ChatGPT (39)

Thesis Table 4.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.8 · source-reported

Silent Failure rate by jurisdiction, aggregated across models and conditions (240 predictions each)

    Jurisdiction         Resource classification        Rate     Approximate count

    Australia            Mid-resource                   30.8%                74
    Nigeria              Low-resource                   23.8%                57
    United States        High-resource                  22.5%                54
    United Kingdom       High-resource                  15.4%                37
Counts are derived from the reported rates over 240 predictions per jurisdiction and sum to the
                                          222 total.

Thesis Table 4.8 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.9 · source-reported

Selected high-confidence incorrect predictions

          Case      Jurisdiction      Stated confidence          Expected Value

          NG 004    Nigeria                    95%                   −0.95
          NG 009    Nigeria                    95%                   −0.95
          US 008    United States             100%                   −1.00
          US 010    United States             100%                   −1.00

Thesis Table 4.9 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 4.10 · source-reported

Cases predicted incorrectly by all six models under No Clue

                                                                  No Clue    With Clue
   Case      Jurisdiction     Legal issue represented
                                                                  correct     correct

   NG 009    Nigeria          Procedural limitation                 0/6         2/6
   NG 020    Nigeria          Document hierarchy                    0/6         5/6
   AU 013    Australia        Purposive statutory construction      0/6         3/6

Thesis Table 4.10 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.