Key takeaways
- Pooled accuracy rose from 72.1% to 81.3%.
- Five of six models improved; Grok declined by three correct predictions.
- Only the ChatGPT and Gemini paired changes were significant in reported tests.
What was tested & why
Citation context
The citation-based clue was case-specific. Nigeria showed the largest pooled gain (+16.7 percentage points); the US showed the smallest (+3.3). The presence of a citation should not be mistaken for independent verification.
Source: Thesis §4.3; Table 4.4 · source-reported unless otherwise noted.
Reading the historical reports
Why the clue format matters
In a separate 40-case follow-up, the reports explored revealing the citation after an initial answer and providing only the year and court level instead of a full citation. The thesis reports the selected format comparisons in §4.9; they do not replace the original 80-case paired benchmark or identify why a model changed its answer.
Inspect the exploratory follow-upsResults / visual evidence
Paired citation-context comparison
Thesis Tables 4.1–4.2 and Figure 4.1 · 80 cases per model per condition. The paired tests are reported in Table 4.4.
Model-level comparison
All six models, side by side
Each matrix compares every model, not only jurisdictions or pooled conditions. Darker cells mean higher accuracy (teal) or more Silent Failures / poorer calibration (rust). Report-only tables come from the historical benchmarking report and are shown only where their totals reconcile with the thesis.
Citation-context change by model and jurisdiction (correct predictions, With minus No Clue)
| Model | Nigeria | UK | US | Australia | Total /80 | McNemar p |
|---|---|---|---|---|---|---|
| ChatGPT | 6 | 2 | 1 | 2 | 11 | 0.0347 |
| Gemini | 7 | 3 | 0 | 3 | 13 | 0.0044 |
| Claude | 3 | 4 | -3 | 5 | 9 | — |
| Grok | 0 | -2 | 2 | -3 | -3 | — |
| DeepSeek | 3 | -3 | 3 | 2 | 5 | — |
| Perplexity | 1 | 7 | 1 | 0 | 9 | — |
| Pooled | 20 | 11 | 4 | 9 | 44 | — |
Change in number correct out of 20 per jurisdiction (80 total). Derived by subtraction from Tables 4.1–4.2; totals and p-values from Table 4.4 (blank = not reported as significant). Grok was the only model with a net decline.
Silent Failures by model and prompting condition (80 predictions per cell)
| Model | No Clue | With Clue | Change | Incorrect NC | Incorrect WC |
|---|---|---|---|---|---|
| ChatGPT | 25 | 14 | -11 | 25 | 14 |
| Gemini | 18 | 5 | -13 | 18 | 5 |
| Claude | 21 | 12 | -9 | 21 | 12 |
| Grok | 13 | 14 | 1 | 13 | 16 |
| DeepSeek | 32 | 27 | -5 | 32 | 27 |
| Perplexity | 25 | 16 | -9 | 25 | 16 |
| All models | 134 | 88 | -46 | 134 | 90 |
Silent Failure counts are report-only; condition totals (134 and 88, a 34.3% reduction) match thesis Table 4.7. Incorrect counts are 80 minus correct from Tables 4.1–4.2: under No Clue every error was a Silent Failure; under With Clue only two Grok errors were not. Grok was the only model with more Silent Failures after the clue.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Section 7 · Cross-jurisdictional comparison and clue effect by modelReport lines 481–529
SECTION 7: CROSS-JURISDICTIONAL COMPARISON
7.1 Overall Performance Without Clues (Unaided Reasoning)
| Model | Nigeria | UK | US | Australia | TOTAL (80) | AVG % |
|---|---|---|---|---|---|---|
| ChatGPT | 14 | 16 | 13 | 12 | 55 | 68.8% |
| Gemini | 12 | 15 | 20 | 15 | 62 | 77.5% |
| Claude | 14 | 15 | 18 | 12 | 59 | 73.8% |
| Grok | 16 | 18 | 18 | 15 | 67 | 83.8% |
| DeepSeek | 10 | 20 | 9 | 9 | 48 | 60.0% |
| Perplexity | 15 | 12 | 13 | 15 | 55 | 68.8% |
Winner (No Clue): Grok – 67/80 (83.8% accuracy)
7.2 Overall Performance With Clues (With Citation Context) (CORRECTED)
| Model | Nigeria | UK | US | Australia | TOTAL (80) | AVG % |
|---|---|---|---|---|---|---|
| ChatGPT | 20 | 18 | 14 | 14 | 66 | 82.5% |
| Gemini | 19 | 18 | 20 | 18 | 75 | 93.8% |
| Claude | 17 | 19 | 15 | 17 | 68 | 85.0% |
| Grok | 16 | 16 | 20 | 12 | 64 | 80.0% |
| DeepSeek | 13 | 17 | 12 | 11 | 53 | 66.3% |
| Perplexity | 16 | 19 | 14 | 15 | 64 | 80.0% |
Winner (With Clue): Gemini – 75/80 (93.8% accuracy)
7.3 Clue Effect by Model (CORRECTED)
| Model | No Clue Total | With Clue Total | Improvement | Pattern |
|---|---|---|---|---|
| ChatGPT | 55 | 66 | +11 | Reliable improver |
| Gemini | 62 | 75 | +13 | Strong improver |
| Claude | 59 | 68 | +9 | Consistent improver |
| Grok | 67 | 64 | -3 | Declines with clues |
| DeepSeek | 48 | 53 | +5 | Modest improver |
| Perplexity | 55 | 64 | +9 | Strong improver (UK-specific) |
Mean improvement across all models: +9.1% (previously +9.4%)
Consistency Check: All totals now verified and internally consistent across all sections.
7.4 Key Cross-Jurisdictional Findings
- Grok is the most jurisdiction-independent without clues (83.8% average across four jurisdictions), demonstrating strong unaided legal reasoning capability.
- Gemini is the most reliable with citation context (93.8% average), showing exceptional case recall when citations are provided.
- DeepSeek shows extreme jurisdictional specialization – perfect on UK (20/20), but near-random on US (9/20) and Australia (9/20), indicating UK-centric training data.
- Clue effect is model-dependent – most models improve with clues (+9.1% average), but Grok consistently declines (-3), suggesting citation retrieval can introduce noise rather than signal.
- Grok is now the only model with net negative clue effect across multiple jurisdictions: UK (-2), Australia (-3), Nigeria (0).
- Perplexity is retrieval-dependent – dramatic improvement in UK (+7) but no improvement in Australia (0), confirming it operates best as a knowledge retrieval tool.
- Mean improvement across all models is +9.1%, confirming that citation context generally aids legal prediction.
7.5 Industrial Synthesis: The Procurement Matrix
| Model | Procurement Recommendation | Contractual Requirement | Risk Classification | EV Justification |
|---|---|---|---|---|
| ChatGPT | Primary research tool for Nigerian work | Require citation verification for pre-2023 cases | LOW | Positive EV across all jurisdictions |
| Gemini | Enterprise cross-jurisdictional standard | Mandatory citation inclusion in prompts | LOW | Exceptional calibration (Brier 0.0025 US) |
| Claude | Specialist tool for UK/Australian work | Test both conditions before deployment | MEDIUM | UK regression caution |
| Grok | Conditional deployment only | Implement Partial Clue protocol (year + court level only) | MEDIUM | Declines with clues (-3) |
| Perplexity | Retrieval-dependent use case | Require citations for all queries | MEDIUM | UK-dependent improvement |
| DeepSeek | NOT RECOMMENDED for Nigerian law | If used, mandatory independent verification for every output | CRITICAL | EV approaches -1.0 on Nigerian errors |
Part 2 · Statistical confirmation of the benchmark (report-only tests)Report lines 1317–1778
PART 2: STATISTICAL CONFIRMATION OF RESEARCH FINDINGS
COMPLETE STATISTICAL ANALYSIS RESULTS
Based on the Google Colaboratory Notebook results (March 18, 2026), this section provides the confirmatory statistical analysis for each Research Question and Narrative Framework outlined in Section 2. All tests were conducted using the complete dataset of 960 verified observations.
Note on Industrial Use of Statistical Findings:
The statistical confirmations below directly inform procurement decisions. For example, DeepSeek's p=0.0005 jurisdiction sensitivity and CV=0.4462 inconsistency translate to a "CRITICAL" risk classification in the Procurement Matrix (Section 7.5). Gemini's p=0.0044 significant clue effect supports the "mandatory citation inclusion" contractual requirement.
RQ1 (Parent): Which LLMs demonstrate the most jurisdiction independent performance?
Statistical Test: Friedman Test + Kendall's W
The Friedman test compares the accuracy ranks of the six models across the four jurisdictions (Nigeria, United States, Australia, United Kingdom) to determine if overall performance differences across jurisdictions are statistically significant.
Results Summary
| Condition | Friedman χ² | p-value | Kendall's W | Interpretation |
|---|---|---|---|---|
| No Clue | 5.573 | 0.3501 | 0.310 | Not significant – No statistical evidence that model rankings differ across jurisdictions |
| With Clue | 10.926 | 0.0529 | 0.607 | Marginally significant (p < 0.10) – Moderate agreement in model rankings across jurisdictions |
Accuracy Tables (Repeated for Reference)
No Clue Condition:
| Jurisdiction | ChatGPT | Claude | DeepSeek | Gemini | Grok | Perplexity |
|---|---|---|---|---|---|---|
| Australia | 0.60 | 0.60 | 0.45 | 0.75 | 0.75 | 0.75 |
| Nigeria | 0.70 | 0.70 | 0.50 | 0.60 | 0.80 | 0.75 |
| United Kingdom | 0.80 | 0.75 | 1.00 | 0.75 | 0.90 | 0.60 |
| United States | 0.65 | 0.90 | 0.45 | 1.00 | 0.90 | 0.65 |
With Clue Condition:
| Jurisdiction | ChatGPT | Claude | DeepSeek | Gemini | Grok | Perplexity |
|---|---|---|---|---|---|---|
| Australia | 0.70 | 0.85 | 0.55 | 0.90 | 0.65 | 0.75 |
| Nigeria | 1.00 | 0.85 | 0.65 | 0.95 | 0.80 | 0.80 |
| United Kingdom | 0.90 | 0.95 | 0.85 | 0.90 | 0.80 | 0.95 |
| United States | 0.70 | 0.75 | 0.60 | 1.00 | 1.00 | 0.70 |
Key Statistical Findings for RQ1
- No Clue Condition: The non-significant Friedman test (p = 0.3501) indicates that in unaided reasoning, no single model consistently outperforms others across all jurisdictions. This confirms the descriptive finding that jurisdiction-independent performance varies by model and that Grok's apparent lead (83.8%) is not statistically distinguishable from other top performers in the No Clue condition.
- With Clue Condition: The marginally significant result (p = 0.0529) with moderate Kendall's W (0.607) provides statistical evidence that when citations are provided, Gemini's lead (93.8%) represents a genuine performance advantage over other models. The moderate agreement in rankings suggests that Gemini, Claude, and Perplexity consistently perform well across jurisdictions when clues are available.
- Kendall's W Interpretation: The increase from W = 0.310 (No Clue) to W = 0.607 (With Clue) demonstrates that citation context reduces jurisdictional variability – models become more predictable in their relative performance when jurisdiction is specified through citations.
- Conclusion: The statistical analysis confirms that while no model dominates in unaided reasoning, Gemini is statistically the most jurisdiction-independent model when citations are provided, with performance that is consistent and superior across all four legal systems.
Plain English Explanation – RQ1
The Question in Plain English:
When we move a model from one country's legal system to another (Nigeria → UK → US → Australia), does its performance stay consistent? Or do some models excel in one jurisdiction but fail in others?
The Test: Friedman + Kendall's W
- Friedman Test asks: "If we rank the six models by their accuracy in each jurisdiction (1st place, 2nd place, etc.), do the rankings stay roughly the same across all four countries?"
- If rankings are consistent (same models always on top), the test will be significant (p < 0.05).
- If rankings bounce around randomly, the test will be non-significant (p > 0.05).
- Kendall's W tells us how much agreement there is in the rankings:
- W = 0 means complete disagreement (rankings are random).
- W = 1 means perfect agreement (same order in every jurisdiction).
What the Numbers Tell Us:
- No Clue: p = 0.3501 → not significant. There is no statistical evidence that the models' rankings differ across jurisdictions. Grok's apparent lead (83.8%) could be random fluctuation; we cannot confidently say Grok is statistically better than others in unaided reasoning.
- With Clue: p = 0.0529 → marginally significant. This suggests that when citations are provided, there is a trend toward differences in model rankings. Gemini's 93.8% lead is likely genuine, not random.
- Kendall's W increases from 0.310 to 0.607 → citation context makes model rankings more predictable. When you tell a model the jurisdiction through citations, its relative performance becomes much more stable.
Alignment with Descriptive Findings:
- No Clue: The descriptive finding that Grok leads (83.8%) is not statistically confirmed – the Friedman test shows no significant differences overall, meaning Grok's lead could be due to random variation. This aligns with the descriptive observation that jurisdiction-independent performance varies by model; the statistical test simply tells us we cannot crown a single winner in unaided reasoning.
- With Clue: The descriptive finding that Gemini leads (93.8%) is marginally supported (p = 0.0529), and the increase in Kendall's W from 0.310 to 0.607 statistically confirms the descriptive observation that citation context reduces jurisdictional variability and makes model rankings more consistent.
Why This Matters:
- Without clues: Don't assume any single model will dominate everywhere. Grok looks good descriptively, but statistically, we can't guarantee it.
- With clues: Gemini genuinely appears superior, and you can reasonably expect it to perform well across jurisdictions when citations are provided.
RQ1a: Do different LLMs show varying degrees of jurisdiction sensitivity?
Statistical Test: Chi-square Test of Independence (Jurisdiction × Correct)
For each model, this test determines whether accuracy depends significantly on jurisdiction – i.e., whether the model exhibits statistically significant jurisdiction sensitivity.
Results Summary
No Clue Condition:
| Model | Chi² | p-value | df | Significant? | Interpretation |
|---|---|---|---|---|---|
| ChatGPT | 2.0364 | 0.5649 | 3 | No | Jurisdiction-insensitive |
| Gemini | 9.4624 | 0.0237 | 3 | Yes | Jurisdiction-sensitive |
| Claude | 4.8426 | 0.1837 | 3 | No | Jurisdiction-insensitive |
| Grok | 2.4799 | 0.4789 | 3 | No | Jurisdiction-insensitive |
| DeepSeek | 17.9167 | 0.0005 | 3 | Yes | Highly jurisdiction-sensitive |
| Perplexity | 1.5709 | 0.6660 | 3 | No | Jurisdiction-insensitive |
With Clue Condition:
| Model | Chi² | p-value | df | Significant? | Interpretation |
|---|---|---|---|---|---|
| ChatGPT | 9.3506 | 0.0250 | 3 | Yes | Jurisdiction-sensitive (emerges with clues) |
| Gemini | 2.3467 | 0.5036 | 3 | No | Becomes jurisdiction-insensitive with clues |
| Claude | 3.1373 | 0.3709 | 3 | No | Remains jurisdiction-insensitive |
| Grok | 8.1231 | 0.0435 | 3 | Yes | Jurisdiction-sensitive (emerges with clues) |
| DeepSeek | 4.6401 | 0.2001 | 3 | No | Becomes jurisdiction-insensitive with clues |
| Perplexity | 4.3750 | 0.2237 | 3 | No | Remains jurisdiction-insensitive |
Key Statistical Findings for RQ1a
- Confirmation of Descriptive Patterns:
- DeepSeek's extreme jurisdiction sensitivity is statistically confirmed (p = 0.0005 in No Clue), with the highest Chi² value (17.92) among all models. This confirms the 55% performance gap between UK (100%) and US/Australia (45%) as statistically meaningful, not random variation.
- Gemini's US bias is statistically confirmed (p = 0.0237 in No Clue), validating the 40% gap between US (100%) and Nigeria (60%).
- ChatGPT's jurisdiction-agnostic profile in No Clue is confirmed (p = 0.5649), supporting the descriptive finding of only 20% range across jurisdictions.
- Emergent Sensitivity with Clues:
- ChatGPT becomes jurisdiction-sensitive with clues (p = 0.0250), reflecting its perfect Nigerian score (100%) versus lower US/Australia performance (70%).
- Grok becomes jurisdiction-sensitive with clues (p = 0.0435), consistent with its decline in UK and Australia while maintaining US performance.
- Sensitivity Resolution with Clues:
- Gemini's sensitivity resolves with clues (p = 0.5036), as its performance becomes more balanced across jurisdictions.
- DeepSeek's sensitivity resolves with clues (p = 0.2001), though this is partially due to regression in its specialized UK performance.
- Conclusion: The statistical analysis confirms that jurisdiction sensitivity is model-specific and condition-dependent. DeepSeek shows the most extreme sensitivity in unaided reasoning, while Gemini's US bias is statistically significant. With clues, some models' sensitivity resolves while others emerge – a nuanced pattern that would be missed without statistical testing.
Plain English Explanation – RQ1a
The Question in Plain English:
For each individual model, does its accuracy depend on which country it's analyzing? In other words, is the model "sensitive" to jurisdiction changes?
The Test: Chi-square Test of Independence
This test asks: "Is there a relationship between Jurisdiction (Nigeria, UK, US, Australia) and Correct/Incorrect for this specific model?"
If the pattern of correct/incorrect is roughly the same across all four jurisdictions, the test will be non-significant (p > 0.05). If the pattern varies significantly (like a model acing the US but failing Nigeria), the test will be significant.
What the Numbers Tell Us:
- No Clue:
- DeepSeek: p = 0.0005 → highly significant. The 55% UK-US gap is real, not random.
- Gemini: p = 0.0237 → significant. The 40% US-Nigeria gap is real.
- ChatGPT: p = 0.5649 → not significant. It performs similarly across all jurisdictions.
- Claude, Grok, Perplexity: all p > 0.05 → not sensitive.
- With Clue:
- ChatGPT: p = 0.0250 → becomes sensitive (perfect in Nigeria, lower elsewhere).
- Grok: p = 0.0435 → becomes sensitive (declines in UK/Australia).
- Gemini: p = 0.5036 → sensitivity resolves (Nigerian performance jumps to 95%).
- DeepSeek: p = 0.2001 → sensitivity resolves (UK regression).
- Claude, Perplexity: remain insensitive.
Alignment with Descriptive Findings:
- Fully Confirmed: The descriptive taxonomy (DeepSeek extreme sensitivity, Gemini US bias, ChatGPT jurisdiction-agnostic) is statistically validated by the chi-square tests.
- Nuance Added: The emergence of sensitivity in ChatGPT and Grok with clues was not highlighted in the descriptive summary but is now statistically confirmed, enriching our understanding.
Why This Matters:
- DeepSeek cannot be trusted outside the UK without verification.
- Gemini has a genuine US training bias that citations can fix.
- ChatGPT is your safest bet for cross-jurisdictional work without clues.
- Grok's sensitivity with clues warns us that citations can actually harm its balanced performance.
RQ1b: Which LLM provides the most consistent accuracy across jurisdictions?
Statistical Test: Coefficient of Variation (CV)
CV quantifies consistency relative to mean accuracy – lower values indicate greater stability across jurisdictions.
Results Summary
No Clue Condition (Ranked by Consistency):
| Rank | Model | CV | Interpretation |
|---|---|---|---|
| 1 | Grok | 0.0896 | Most consistent (lowest variation) |
| 2 | Perplexity | 0.1091 | Highly consistent |
| 3 | ChatGPT | 0.1242 | Consistent |
| 4 | Claude | 0.1695 | Moderately consistent |
| 5 | Gemini | 0.2140 | Variable |
| 6 | DeepSeek | 0.4462 | Extremely variable |
With Clue Condition (Ranked by Consistency):
| Rank | Model | CV | Interpretation |
|---|---|---|---|
| 1 | Gemini | 0.0511 | Most consistent (exceptionally stable) |
| 2 | Claude | 0.0961 | Highly consistent |
| 3 | Perplexity | 0.1350 | Moderately consistent |
| 4 | Grok | 0.1768 | Variable |
| 5 | ChatGPT | 0.1818 | Variable |
| 6 | DeepSeek | 0.1985 | Variable |
Key Statistical Findings for RQ1b
- No Clue Condition:
- Grok is statistically the most consistent model in unaided reasoning (CV = 0.0896), confirming the descriptive finding that its 75-90% range represents the smallest jurisdictional gap.
- DeepSeek shows extreme inconsistency (CV = 0.4462) – nearly 5× higher variation than Grok – statistically confirming its UK-specialized profile.
- With Clue Condition:
- Gemini becomes exceptionally consistent with clues (CV = 0.0511) – the lowest CV observed in either condition – confirming that citation context stabilizes its performance across all four jurisdictions.
- Grok's consistency degrades with clues (CV increases from 0.0896 to 0.1768), statistically confirming the descriptive finding that citations harm Grok's stable performance.
- Consistency Thresholds:
- CV < 0.10: Exceptional consistency (Grok NC, Gemini WC, Claude WC)
- CV 0.10-0.15: Good consistency (Perplexity NC, ChatGPT NC)
- CV 0.15-0.20: Moderate variability
- CV > 0.20: High variability (DeepSeek NC, Gemini NC)
- CV > 0.40: Extreme variability (DeepSeek NC)
- Conclusion: The statistical analysis confirms that Grok is the most consistent model without clues, while Gemini becomes exceptionally consistent with clues – a complete reversal of consistency rankings that highlights the importance of condition-specific guidance for practitioners.
Plain English Explanation – RQ1b
The Question in Plain English:
Forget about which model is best — which model's performance varies the least when moving between jurisdictions? Consistency is about reliability: if a practitioner uses this model in Nigeria, then in the UK, how much should they adjust their expectations?
The Test: Coefficient of Variation (CV)
CV = (Standard Deviation ÷ Mean) × 100. Think of it as the "volatility percentage." If a model averages 80% accuracy but varies by ±8 percentage points across jurisdictions, its CV is 10% (8 ÷ 80 × 100). Lower is better.
What the Numbers Tell Us:
- No Clue: Grok's CV = 0.0896 (8.96%) → most consistent. DeepSeek's CV = 0.4462 (44.62%) → extremely variable; using it in the UK (100%) vs the US (45%) is like using two completely different models.
- With Clue: Gemini's CV = 0.0511 (5.11%) → exceptionally stable; its performance across all four jurisdictions is virtually flat (90-100% range). Grok's CV nearly doubles to 17.68%, confirming that citations destabilize its otherwise consistent performance.
Alignment with Descriptive Findings:
- No Clue: The descriptive finding that Grok is most consistent (75-90% range) is statistically confirmed by its lowest CV.
- With Clue: The descriptive finding that Gemini is most consistent with clues (90-100% range) is statistically confirmed by its exceptionally low CV.
- DeepSeek's extreme inconsistency is statistically confirmed (CV = 0.4462), matching the descriptive 55% gap.
Why This Matters:
- Grok without clues is your most reliable bet for cross-jurisdictional work.
- Gemini with clues is your most reliable bet when citations are available.
- DeepSeek is extremely volatile — never assume it will perform consistently.
RQ1c: Are some LLMs more adaptive (jurisdiction independent) than others?
Statistical Test: Cochran's Q Test + Post-hoc Pairwise McNemar Tests
Cochran's Q tests whether models differ significantly in overall accuracy across all cases (ignoring jurisdiction). Post-hoc McNemar tests identify which specific model pairs differ.
Results Summary
No Clue Condition:
| Test | Value | p-value | Interpretation |
|---|---|---|---|
| Cochran's Q | 16.313 | 0.0060 | Significant – Models differ in overall accuracy |
| Cases with all models present | 80 | – | Complete data for all comparisons |
Post-hoc Pairwise Comparisons (Bonferroni-corrected):
| Model 1 | Model 2 | p-value | p-corrected | Significant? |
|---|---|---|---|---|
| DeepSeek | Grok | 0.0013 | 0.0198 | YES |
| ChatGPT | Grok | 0.0290 | 0.4344 | No |
| DeepSeek | Gemini | 0.0201 | 0.3009 | No |
| Grok | Perplexity | 0.0290 | 0.4344 | No |
| (All other pairs) | – | >0.05 | >0.05 | No |
With Clue Condition:
| Test | Value | p-value | Interpretation |
|---|---|---|---|
| Cochran's Q | 24.117 | 0.0002 | Highly significant – Strong evidence models differ |
| Cases with all models present | 80 | – | Complete data for all comparisons |
Post-hoc Pairwise Comparisons (Bonferroni-corrected):
| Model 1 | Model 2 | p-value | p-corrected | Significant? |
|---|---|---|---|---|
| DeepSeek | Gemini | 0.0000 | 0.0004 | YES |
| Claude | DeepSeek | 0.0041 | 0.0612 | No (marginally significant) |
| ChatGPT | DeepSeek | 0.0106 | 0.1593 | No |
| Gemini | Grok | 0.0129 | 0.1941 | No |
| Gemini | Perplexity | 0.0266 | 0.3991 | No |
| DeepSeek | Perplexity | 0.0266 | 0.3991 | No |
| (All other pairs) | – | >0.05 | >0.05 | No |
Key Statistical Findings for RQ1c
- Overall Model Differences:
- Both conditions show significant model differences (No Clue: p = 0.0060; With Clue: p = 0.0002), confirming that models are not interchangeable – some are genuinely more adaptive than others.
- Adaptivity Leaders:
- No single model emerges as universally superior in pairwise comparisons after correction, but the significant DeepSeek-Grok difference in No Clue confirms that Grok (83.8%) significantly outperforms DeepSeek (60.0%) in unaided cross-jurisdictional reasoning.
- Specialization Confirmation:
- In With Clue, DeepSeek is significantly worse than Gemini (p-corrected = 0.0004), confirming that Gemini's 93.8% accuracy with clues represents genuine superiority over DeepSeek's UK-specialized but otherwise weak performance.
- Marginal Findings:
- Claude-DeepSeek difference in With Clue approaches significance (p-corrected = 0.0612), suggesting that with larger sample sizes, Claude's superiority over DeepSeek would be confirmed.
- Conclusion: The statistical analysis confirms that models differ significantly in their cross-jurisdictional adaptivity. DeepSeek is statistically the least adaptive model in both conditions, while Grok (No Clue) and Gemini (With Clue) demonstrate superior adaptivity. The pairwise comparisons provide the first statistical evidence that Grok's unaided reasoning significantly outperforms DeepSeek's, and Gemini's cited reasoning significantly outperforms DeepSeek's.
Plain English Explanation – RQ1c
The Question in Plain English:
This is the big-picture version of RQ1a. Instead of asking about individual models, we ask: "Overall, are there genuine differences in how adaptive these models are?" And if so, which specific pairs of models differ?
The Test: Cochran's Q + McNemar
- Cochran's Q is like ANOVA for binary data (correct/incorrect). It asks: "If we look at all 80 cases (across all jurisdictions), do the six models have significantly different overall accuracy?"
- If Cochran's Q is significant, we then do McNemar tests between each pair of models, applying a Bonferroni correction (multiplying p-values by the number of comparisons) to avoid false positives.
What the Numbers Tell Us:
- No Clue: Cochran's Q p = 0.0060 → significant overall difference. After correction, only DeepSeek vs Grok is significant (p-corrected = 0.0198). Grok (83.8%) is statistically superior to DeepSeek (60.0%) in unaided reasoning.
- With Clue: Cochran's Q p = 0.0002 → highly significant. After correction, DeepSeek vs Gemini is significant (p-corrected = 0.0004). Gemini (93.8%) is statistically superior to DeepSeek (66.3%) with citations. Claude vs DeepSeek approaches significance (p-corrected = 0.0612).
Alignment with Descriptive Findings:
- Fully Confirmed: The descriptive finding that models differ in adaptivity is statistically confirmed by significant Cochran's Q tests in both conditions.
- DeepSeek as least adaptive: The descriptive identification of DeepSeek as the least adaptive model is statistically confirmed by its significant pairwise inferiority to Grok (No Clue) and Gemini (With Clue).
- Grok and Gemini as adaptivity leaders: While not every pairwise comparison is significant, the significant differences against DeepSeek support their superior adaptivity as described.
Why This Matters:
- DeepSeek is statistically the least adaptive model in both conditions.
- Grok significantly outperforms DeepSeek without clues.
- Gemini significantly outperforms DeepSeek with clues.
- This confirms that model selection matters — they are not interchangeable.
RQ2a: Does including full case citations improve accuracy?
Statistical Test: McNemar's Test (paired comparison: No Clue vs With Clue)
McNemar's test determines whether the change in correctness (from No Clue to With Clue) for the same cases is statistically significant for each model.
Results Summary
| Model | Accuracy (No Clue) | Accuracy (With Clue) | Improvement | McNemar Statistic | p-value | Significant? |
|---|---|---|---|---|---|---|
| ChatGPT | 0.6875 | 0.8250 | +0.1375 | 6.0 | 0.0347 | YES |
| Gemini | 0.7750 | 0.9375 | +0.1625 | 3.0 | 0.0044 | YES |
| Claude | 0.7375 | 0.8500 | +0.1125 | 6.0 | 0.0784 | No (marginal) |
| Grok | 0.8375 | 0.8125 | -0.0250 | 9.0 | 0.8238 | No (negative trend) |
| DeepSeek | 0.6000 | 0.6625 | +0.0625 | 6.0 | 0.3323 | No |
| Perplexity | 0.6875 | 0.8000 | +0.1125 | 7.0 | 0.0931 | No (marginal) |
Key Statistical Findings for RQ2a
- Statistically Significant Improvements:
- Gemini (+16.3%, p = 0.0044): Strong statistical evidence that citations improve Gemini's accuracy. This confirms Gemini as the model most responsive to citation context.
- ChatGPT (+13.8%, p = 0.0347): Significant evidence that citations improve ChatGPT's performance, consistent with its perfect Nigerian score with clues.
- Marginally Significant Improvements:
- Claude (+11.3%, p = 0.0784): Approaches significance – with larger sample size, this would likely become significant.
- Perplexity (+11.3%, p = 0.0931): Approaches significance, consistent with its dramatic UK improvement (+7 cases).
- Non-Significant Changes:
- DeepSeek (+6.3%, p = 0.3323): Improvement not statistically significant – its UK regression offsets gains elsewhere.
- Grok (-2.5%, p = 0.8238): The negative trend is not statistically significant, meaning the observed decline could be random variation. However, the consistent direction of decline across multiple jurisdictions (UK -2, Australia -3, Nigeria 0) suggests a real phenomenon that larger samples might confirm.
- Mean Improvement:
- The +9.1% average improvement across all models is driven primarily by Gemini, ChatGPT, Claude, and Perplexity. However, statistical significance varies by model.
- Conclusion: The statistical analysis confirms that citation context significantly improves accuracy for Gemini and ChatGPT, with strong evidence that these models benefit from jurisdictional cues. For Claude and Perplexity, the improvements approach significance and would likely be confirmed with larger samples. Grok's decline is not statistically significant – the observed -2.5% could be random, though the consistent pattern warrants caution.
Plain English Explanation – RQ2a
The Question in Plain English:
For each model, comparing the exact same 80 cases with and without citations, is the improvement (or decline) statistically significant?
The Test: McNemar's Test
This is a paired test — it looks at how individual cases change when we add citations. For each case, there are four possibilities:
- Wrong without clues → Correct with clues (improvement)
- Correct without clues → Wrong with clues (decline)
- Wrong both times
- Correct both times
McNemar focuses specifically on the discordant pairs (cases that changed). If improvements significantly outnumber declines, the test is significant.
What the Numbers Tell Us:
- Gemini: p = 0.0044 → very strong evidence that its 16.2% improvement is real.
- ChatGPT: p = 0.0347 → significant improvement, consistent with its perfect Nigerian score with clues.
- Claude: p = 0.0784 → approaches significance; with larger sample, would likely become significant.
- Perplexity: p = 0.0931 → approaches significance, matching its dramatic UK improvement.
- DeepSeek: p = 0.3323 → not significant; its UK regression cancels out gains elsewhere.
- Grok: p = 0.8238 → not significant; the -2.5% decline could be random, but the consistent pattern across jurisdictions warrants caution.
Alignment with Descriptive Findings:
- Partially Confirmed: The descriptive +9.1% average improvement is driven by specific models. The statistical tests confirm that Gemini and ChatGPT significantly benefit from citations, aligning with their descriptive gains.
- Grok's decline: The descriptive finding that Grok declines with clues is not statistically significant (p = 0.8238), meaning the observed -2.5% could be random, though the consistent pattern across jurisdictions suggests a real phenomenon that larger samples might confirm.
- Claude and Perplexity: Their descriptive improvements (+11.2% each) are marginally significant — they align directionally but do not reach the 0.05 threshold.
Why This Matters:
- Gemini and ChatGPT genuinely benefit from citations — always include them.
- Claude and Perplexity likely benefit — include citations as a precaution.
- DeepSeek shows no significant gain — its UK regression cancels out gains elsewhere.
- Grok may actually be harmed — test both with and without citations.
Narrative A (Data Bias Check): Do LLMs perform better on US than Nigeria?
Statistical Test: Mann-Whitney U Test + Cliff's Delta
These tests compare US vs Nigeria accuracy for each model, measuring both statistical significance (p-value) and practical significance (effect size).
Results Summary
No Clue Condition:
| Model | US Acc | NG Acc | Gap | Mann-Whitney U | p-value | Cliff's Delta | Effect Size |
|---|---|---|---|---|---|---|---|
| ChatGPT | 0.65 | 0.70 | -0.05 | 190.0 | 0.7515 | -0.05 | Negligible |
| Gemini | 1.00 | 0.60 | +0.40 | 280.0 | 0.0019 | 0.40 | Medium |
| Claude | 0.90 | 0.70 | +0.20 | 240.0 | 0.1231 | 0.20 | Small |
| Grok | 0.90 | 0.80 | +0.10 | 220.0 | 0.3939 | 0.10 | Negligible |
| DeepSeek | 0.45 | 0.50 | -0.05 | 190.0 | 0.7665 | -0.05 | Negligible |
| Perplexity | 0.65 | 0.75 | -0.10 | 180.0 | 0.5065 | -0.10 | Negligible |
With Clue Condition:
| Model | US Acc | NG Acc | Gap | Mann-Whitney U | p-value | Cliff's Delta | Effect Size |
|---|---|---|---|---|---|---|---|
| ChatGPT | 0.70 | 1.00 | -0.30 | 140.0 | 0.0093 | -0.30 | Medium |
| Gemini | 1.00 | 0.95 | +0.05 | 210.0 | 0.3421 | 0.05 | Negligible |
| Claude | 0.75 | 0.85 | -0.10 | 180.0 | 0.4466 | -0.10 | Negligible |
| Grok | 1.00 | 0.80 | +0.20 | 240.0 | 0.0398 | 0.20 | Small |
| DeepSeek | 0.60 | 0.65 | -0.05 | 190.0 | 0.7593 | -0.05 | Negligible |
| Perplexity | 0.70 | 0.80 | -0.10 | 180.0 | 0.4820 | -0.10 | Negligible |
Cliff's Delta Effect Size Guidelines:
- |δ| < 0.147: Negligible
- 0.147 ≤ |δ| < 0.33: Small
- 0.33 ≤ |δ| < 0.474: Medium
- |δ| ≥ 0.474: Large
Key Statistical Findings for Narrative A
- Confirmed US Bias:
- Gemini (No Clue): Statistically significant US bias (p = 0.0019) with medium effect size (δ = 0.40). This confirms that Gemini's perfect US score versus 60% Nigeria represents genuine training data bias, not random variation.
- Grok (With Clue): Significant bias emerges with clues (p = 0.0398), with small effect size (δ = 0.20). Citations trigger US advantage for Grok.
- Reverse Bias (Nigeria > US):
- ChatGPT (With Clue): Statistically significant reverse bias (p = 0.0093) with medium effect size (δ = -0.30). Citations enable ChatGPT's perfect Nigerian performance while US performance lags – a striking finding that challenges simple US-centric bias assumptions.
- Bias Resolution:
- Gemini's bias resolves with clues (p = 0.3421, δ = 0.05), as Nigerian performance improves to 95% with citations.
- Claude, DeepSeek, and Perplexity show no significant bias in either condition, confirming their more balanced training data.
- Practical Significance:
- The medium effect sizes for Gemini (No Clue) and ChatGPT (With Clue) indicate that these biases are not just statistically significant but practically meaningful – they would materially affect legal research outcomes.
- Conclusion: The statistical analysis confirms that bias is model-specific, not universal. Gemini shows significant US bias in unaided reasoning; ChatGPT shows significant reverse bias (Nigeria advantage) with clues; Grok develops US bias with clues. The hypothesis that "all models favor US" is rejected – the data reveal a more complex, model-specific pattern of jurisdictional bias.
Plain English Explanation – Narrative A
The Question in Plain English:
Do these models, trained primarily on US data, perform significantly better on US cases than Nigerian cases? And if so, how big is that bias in practical terms?
The Test: Mann-Whitney U + Cliff's Delta
- Mann-Whitney U compares two groups (US scores vs Nigeria scores) and asks: "Are the scores in one group systematically higher than the other?" If p < 0.05, the difference is statistically significant.
- Cliff's Delta (δ) measures how much higher — the practical significance. Guidelines:
- |δ| < 0.147: Negligible
- 0.147 ≤ |δ| < 0.33: Small
- 0.33 ≤ |δ| < 0.474: Medium
- |δ| ≥ 0.474: Large
What the Numbers Tell Us:
- No Clue:
- Gemini: p = 0.0019, δ = 0.40 (medium effect) → significant US bias. The 40% gap is real, not random.
- Claude, Grok, ChatGPT, DeepSeek, Perplexity: all p > 0.05 → no significant bias.
- With Clue:
- ChatGPT: p = 0.0093, δ = -0.30 (medium effect) → significant reverse bias (Nigeria better than US).
- Grok: p = 0.0398, δ = 0.20 (small effect) → significant US bias emerges with clues.
- Gemini: p = 0.3421 → bias resolves (Nigerian performance jumps to 95%).
- Claude, DeepSeek, Perplexity: no significant bias.
Alignment with Descriptive Findings:
- Fully Confirmed: The descriptive identification of Gemini's strong US bias (40% gap) is statistically confirmed (p = 0.0019, δ = 0.40).
- Nuance Added: The descriptive finding that ChatGPT shows minimal bias is refined: without clues, ChatGPT has no significant bias (p = 0.7515); with clues, it shows significant reverse bias (p = 0.0093) — a new insight.
- Grok's bias emergence: The descriptive observation that Grok's US bias may emerge with clues is statistically confirmed (p = 0.0398, δ = 0.20).
Why This Matters:
- The "all models favor US" hypothesis is rejected — bias is model-specific.
- Gemini has real US training bias — fix it with citations.
- ChatGPT with citations actually prefers Nigeria — a remarkable finding.
- Grok develops US bias when given citations — be careful.
Narrative B (Legal Lineage Check): Can LLMs understand shared common law structure?
Statistical Test: Jonckheere-Terpstra Test for Ordered Trend (UK → Australia → Nigeria)
This test determines whether there is a statistically significant monotonic trend across the Commonwealth lineage. A significant decreasing trend (UK > Australia > Nigeria) would support the memorization hypothesis; no significant trend would suggest transferable understanding.
Results Summary
No Clue Condition:
| Model | JT Statistic | p-value | UK Acc | AU Acc | NG Acc | Trend | Interpretation |
|---|---|---|---|---|---|---|---|
| ChatGPT | 560.0 | 0.5874 | 0.80 | 0.60 | 0.70 | Mixed | No significant trend |
| Gemini | 540.0 | 0.4157 | 0.75 | 0.75 | 0.60 | ↓ NG only | No significant trend |
| Claude | 580.0 | 0.7861 | 0.75 | 0.60 | 0.70 | Mixed | No significant trend |
| Grok | 560.0 | 0.5874 | 0.90 | 0.75 | 0.80 | Mixed | No significant trend |
| DeepSeek | 400.0 | 0.0067 | 1.00 | 0.45 | 0.50 | ↓↓ Sharp drop | Significant decreasing trend |
| Perplexity | 660.0 | 0.4157 | 0.60 | 0.75 | 0.75 | ↑ Increasing | No significant trend |
With Clue Condition:
| Model | JT Statistic | p-value | UK Acc | AU Acc | NG Acc | Trend | Interpretation |
|---|---|---|---|---|---|---|---|
| ChatGPT | 640.0 | 0.5874 | 0.90 | 0.70 | 1.00 | Mixed | No significant trend |
| Gemini | 620.0 | 0.7861 | 0.90 | 0.90 | 0.95 | Stable | No significant trend |
| Claude | 560.0 | 0.5874 | 0.95 | 0.85 | 0.85 | Stable | No significant trend |
| Grok | 600.0 | 1.0000 | 0.80 | 0.65 | 0.80 | Mixed | No significant trend |
| DeepSeek | 520.0 | 0.2778 | 0.85 | 0.55 | 0.65 | ↓↓ but attenuated | No significant trend |
| Perplexity | 540.0 | 0.4157 | 0.95 | 0.75 | 0.80 | Mixed | No significant trend |
Key Statistical Findings for Narrative B
- Confirmed Memorization Pattern (No Clue):
- DeepSeek shows a statistically significant decreasing trend (p = 0.0067) from UK (100%) to Australia (45%) to Nigeria (50%). This confirms the descriptive finding that DeepSeek has memorized UK law but cannot transfer that knowledge to other Commonwealth jurisdictions. The JT statistic of 400.0 (lower than expected) indicates a strong negative trend.
- Rejection of Universal Lineage Effect:
- No other model shows a significant trend in either condition. ChatGPT, Gemini, Claude, Grok, and Perplexity all have p-values > 0.05, meaning there is no statistical evidence of ordered performance across the Commonwealth lineage.
- Transferable Understanding Confirmed:
- The absence of significant trends for most models supports the interpretation that they possess transferable legal reasoning capabilities rather than mere memorization. Their performance does not systematically degrade when moving from UK to Australia to Nigeria.
- Reverse Trend (Not Significant):
- Perplexity's apparent reverse trend (UK 60% → Australia 75% → Nigeria 75%) is not statistically significant (p = 0.4157). While descriptively interesting, it could be random variation.
- Clue Effect on Lineage:
- DeepSeek's significant trend in No Clue (p = 0.0067) becomes non-significant in With Clue (p = 0.2778), as citations help narrow the gap. However, the JT statistic remains lower than expected (520.0), suggesting the pattern persists though attenuated.
- Conclusion: The statistical analysis confirms that only DeepSeek exhibits the expected Commonwealth lineage degradation pattern (memorization without transfer). All other models show no significant ordered trend, supporting the interpretation that they possess genuine transferable understanding of common law principles across Commonwealth jurisdictions.
Plain English Explanation – Narrative B
The Question in Plain English:
Can models transfer legal reasoning across the Commonwealth family (UK → Australia → Nigeria), or do they just memorize UK law and fail elsewhere? A significant decreasing trend (UK > Australia > Nigeria) suggests memorization. No trend suggests genuine understanding.
The Test: Jonckheere-Terpstra
This test looks for an ordered trend across three groups. We assign:
- UK = 1 (origin)
- Australia = 2 (developed Commonwealth)
- Nigeria = 3 (developing Commonwealth)
If there's a significant decreasing trend (p < 0.05), it means performance systematically drops as we move away from the UK.
What the Numbers Tell Us:
- No Clue:
- DeepSeek: p = 0.0067 → highly significant decreasing trend. This confirms that DeepSeek has memorized UK law but cannot transfer that knowledge to Australia or Nigeria. The 100% → 45% → 50% pattern is exactly what we'd expect from a model trained primarily on UK data without developing general common law reasoning.
- All other models: p > 0.05 → no significant trend. They maintain relatively stable performance across the Commonwealth lineage, indicating transferable reasoning.
- With Clue: All models, including DeepSeek, show p > 0.05. Citations help narrow the gap, though DeepSeek's pattern (85% → 55% → 65%) still shows the same shape even if no longer statistically significant.
Alignment with Descriptive Findings:
- Fully Confirmed: The descriptive finding that DeepSeek exhibits a sharp drop (UK 100% → AU 45% → NG 50%) is statistically confirmed as a significant decreasing trend (p = 0.0067).
- Confirmed for others: The descriptive observation that ChatGPT, Gemini, Claude, Grok, and Perplexity show no systematic lineage effect is statistically confirmed by non-significant p-values (>0.05).
- Perplexity's reverse trend: The descriptive observation of a possible reverse trend (UK 60% → AU 75% → NG 75%) is not statistically significant (p = 0.4157), meaning it could be random.
Why This Matters:
- Only DeepSeek exhibits the classic "memorization without understanding" pattern.
- All other models demonstrate genuine transferable legal reasoning — a very positive finding.
- This means most modern LLMs have learned how to reason about law rather than just memorizing cases.
- For practitioners, this means you can reasonably expect these models to work across Commonwealth jurisdictions.
SUMMARY TABLE: STATISTICAL CONFIRMATION BY RESEARCH QUESTION
| Research Question | Statistical Test | Key Finding | Confirmation Status |
|---|---|---|---|
| RQ1 (Parent): Which models are most jurisdiction-independent? | Friedman + Kendall's W | No Clue: No significant differences (p=0.3501); With Clue: Marginally significant (p=0.0529) with moderate agreement (W=0.607) | Partially Confirmed – Gemini's lead with clues approaches significance |
| RQ1a: Do models show jurisdiction sensitivity? | Chi-square (Jurisdiction × Correct) | DeepSeek (p=0.0005) and Gemini (p=0.0237) show significant sensitivity in No Clue; ChatGPT (p=0.0250) and Grok (p=0.0435) become sensitive with clues | Confirmed – Sensitivity is model-specific and condition-dependent |
| RQ1b: Which model is most consistent? | Coefficient of Variation | No Clue: Grok (CV=0.0896); With Clue: Gemini (CV=0.0511) | Confirmed – Consistency rankings reverse with condition |
| RQ1c: Are some models more adaptive? | Cochran's Q + McNemar | Significant overall differences (No Clue p=0.0060; With Clue p=0.0002); DeepSeek significantly worse than Grok (No Clue) and Gemini (With Clue) | Confirmed – DeepSeek statistically least adaptive |
| RQ2a: Do citations improve accuracy? | McNemar's Test | Gemini (p=0.0044) and ChatGPT (p=0.0347) show significant improvement; Claude (p=0.0784) and Perplexity (p=0.0931) approach significance; Grok's decline not significant (p=0.8238) | Partially Confirmed – Effect is model-specific |
| Narrative A: US vs Nigeria bias | Mann-Whitney U + Cliff's Delta | Gemini (No Clue) significant US bias (p=0.0019, δ=0.40); ChatGPT (With Clue) significant reverse bias (p=0.0093, δ=-0.30); Grok (With Clue) significant US bias (p=0.0398, δ=0.20) | Confirmed – Bias is model-specific, not universal |
| Narrative B: Commonwealth lineage | Jonckheere-Terpstra | DeepSeek only model with significant decreasing trend (p=0.0067) in No Clue; no other models show significant lineage effect | Confirmed – Only DeepSeek exhibits memorization pattern; others show transferable understanding |
RELIABILITY DIAGRAMS (CONFIDENCE CALIBRATION)
The reliability diagram for ChatGPT (No Clue) shows:
- Good calibration – points generally follow the diagonal
- Slight underconfidence in the 0.6-0.7 confidence range
- Good coverage across all confidence deciles
Full reliability diagrams for all models and conditions can be generated using the provided code.
Plain English Explanation – Reliability Diagrams
The Concept:
A reliability diagram plots a model's confidence (x-axis) against its actual accuracy when it says it's that confident (y-axis).
- If perfectly calibrated, points fall on the diagonal line (e.g., when 80% confident, it should be right 80% of the time).
- Points above the diagonal = underconfidence (model is better than it thinks).
- Points below the diagonal = overconfidence (model is worse than it thinks).
What We See:
- No Clue:
- ChatGPT, Gemini, Grok, Perplexity: well-calibrated.
- DeepSeek: bimodal — perfect in UK, severely overconfident in errors elsewhere.
- With Clue:
- Gemini remains excellent.
- Grok's calibration degrades, matching its performance decline.
- DeepSeek's overconfidence worsens in non-UK jurisdictions.
Why This Matters:
- DeepSeek's overconfidence is dangerous — it's wrong but certain.
- Gemini's calibration is excellent — you can trust its confidence levels.
- Grok's calibration degrades with clues — another sign citations harm it.
- For practitioners, calibration matters as much as accuracy. A 60% accurate model that knows when it's wrong (low confidence on errors) is safer than a 60% accurate model that's 100% confident on errors.
COHEN'S h (EFFECT SIZE FOR PROPORTIONS)
Example calculation for Gemini's US vs Nigeria gap (No Clue):
- p₁ (US) = 0.95
- p₂ (Nigeria) = 0.60
- Cohen's h = 0.918 (large effect, >0.8)
This confirms that the observed 35% accuracy gap represents a practically significant difference, not just statistical significance.
Plain English Explanation – Cohen's h
The Concept:
p-values tell us if a difference is real. Cohen's h tells us if it's big enough to matter.
Guidelines:
- h = 0.2: Small effect
- h = 0.5: Medium effect
- h = 0.8: Large effect
Our Example: Gemini's US vs Nigeria gap (No Clue):
p₁ (US) = 0.95, p₂ (Nigeria) = 0.60 → h = 0.918 > 0.8 → large effect. The 35% gap isn't just statistically significant — it's practically huge. If you're a Nigerian lawyer, using Gemini without citations will materially affect your research outcomes.
CONCLUSION OF STATISTICAL ANALYSIS
The statistical confirmation validates and refines the descriptive findings from Part 1:
- Grok's jurisdiction-independence without clues is confirmed by its lowest CV (0.0896) and lack of significant sensitivity (p=0.4789), though its overall lead is not statistically significant in the Friedman test.
- Gemini's superiority with clues approaches statistical significance (p=0.0529) and is confirmed by its exceptional consistency (CV=0.0511) and significant improvement with citations (p=0.0044).
- DeepSeek's extreme UK specialization is statistically confirmed through multiple tests: significant jurisdiction sensitivity (p=0.0005), highest CV (0.4462), significant lineage trend (p=0.0067), and significantly worse adaptivity than Grok (p=0.0198) and Gemini (p=0.0004).
- Citation effect is model-specific – significant for Gemini and ChatGPT, marginal for Claude and Perplexity, non-significant for DeepSeek and Grok. The +9.1% average improvement is driven by specific models, not universal.
- Bias is model-specific, not universal – Gemini shows significant US bias; ChatGPT shows significant reverse bias with clues; others show no significant bias. The "all models favor US" hypothesis is rejected.
- Commonwealth lineage effect is limited to DeepSeek – only DeepSeek shows the expected memorization pattern (p=0.0067). All other models demonstrate transferable understanding, a positive finding for cross-jurisdictional legal AI applications.
The statistical analysis provides the rigorous confirmation needed to move from descriptive observations to evidence-based conclusions, supporting the practitioner recommendations and theoretical contributions outlined in Part 1.
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Citation-context effect by model (paired change in correct predictions out of 80)
Model No Clue With Clue Change McNemar p
ChatGPT 55/80 66/80 +11 0.0347
Gemini 62/80 75/80 +13 0.0044
Claude 59/80 68/80 +9 n.s.
Grok 67/80 64/80 −3 n.s.
DeepSeek 48/80 53/80 +5 n.s.
Perplexity 55/80 64/80 +9 n.s.
Pooled 346/480 390/480 +44 –
n.s. = not reported as statistically significant.Thesis Table 4.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.