Key takeaways
- EQ1 explored Grok clue format.
- EQ2 checked DeepSeek UK citation reproduction and reran cases.
- EQ3 tested ChatGPT prompt sensitivity on selected errors.
What was tested & why
Continuation experiments
These experiments are supplementary exploratory analyses. DeepSeek’s 20/20 UK No Clue result fell to 16/20 on rerun. Citation reproduction cannot confirm memorisation, and selected-case prompt changes do not establish a universal mechanism.
Supplementary exploratory analysis; selected cases and formats limit inference.
Source: Thesis §§3.11, 4.9–4.10; Table 3.7 · source-reported unless otherwise noted.
Historical reports · supplementary exploratory analysis
What the earlier experiment logs add
EQ1 · Grok clue formats
The weekly report describes five formats: No Clue, full citation, Late Reveal after an initial verdict, Partial Clue containing year and court level, and Jurisdiction Only with year. The 40-case follow-up table is reproduced below. These are format-specific observations, not the primary With Clue results.
| Condition | Australia correct | UK correct |
|---|---|---|
| No Clue | 15/20 · 75% | 18/20 · 90% |
| Full citation | 13/20 · 65% | 16/20 · 80% |
| Late Reveal | 12/20 · 60% | 19/20 · 95% |
| Partial Clue | 18/20 · 90% | 17/20 · 85% |
| Jurisdiction Only | 17/20 · 85% | 15/20 · 75% |
Thesis §4.9 confirms the 90% Australian Partial Clue, 95% UK Late Reveal, 60% Australian Late Reveal and 65% Australian full-citation comparisons. The rest of this table remains report-only.
EQ2 · DeepSeek rerun
Thesis §4.9 reports citation reproduction from sanitised facts in at least 40% of examined UK cases and a fall from 20/20 to 16/20 on a fresh UK rerun. The reports explore case familiarity, but neither the reproduction nor the rerun proves memorisation or reveals training data.
EQ3 · ChatGPT selected errors
In the weekly report, 8 of 15 initial No-Clue errors changed on fresh rerun; the other seven changed after adversarial pushback. Jurisdiction Masking changed 2/8 selected Australian errors, while Confidence Penalty changed 3/6 selected Australian With-Clue errors. Thesis §4.9 consolidates the prompt-sensitivity result but does not include these subset counts. They are historical report-only, conditional on selected errors, and do not establish a general correction rate or mechanism.
Weekly benchmarking report, EQ1–EQ3 follow-up passages; final thesis §4.9 takes precedence. Counts beyond the consolidated thesis record are report-only, not independently re-estimated. No prompt is recommended as a reliable correction technique.
Benchmarking report · report-only · full record
What the benchmarking report adds to this page
These sections are reproduced in full from the Complete Research Findings Report (Part I benchmark, March–April 2026), the earlier working record behind this part of the study. They are report-only: the thesis is the authority wherever numbers or interpretations differ. Known differences: the report’s +9.1% (mean of model-level gains) against the thesis’s pooled +9.2 points; its significance tests, logistic regression and power analysis, which the thesis does not report; wording such as “bias”, “proves” or “memorisation”, which the thesis avoids because these results cannot establish mechanisms; and its timeline dates. Reasoning grades are the researcher’s own review and are still awaiting lawyer verification.
Part 3 · From statistical confirmation to follow-up experiment designReport lines 1779–1998
PART 3 — NEXT STEPS: FROM STATISTICAL CONFIRMATION TO MECHANISTIC UNDERSTANDING
The statistical analysis has confirmed that certain phenomena exist, but it cannot explain why they occur. This section outlines the next phase of research designed to uncover the mechanisms behind our most robust findings. Each experiment builds directly on the statistical evidence already established.
What Statistical Analysis Has Already Told Us About These Phenomena
Before detailing the experiments, we summarize what the statistical analysis has already revealed about each target phenomenon:
| Phenomenon | Statistical Evidence | Status |
|---|---|---|
| Grok's decline with clues | Pattern: UK -2, AU -3, NG 0; McNemar p = 0.8238 | Pattern consistent but not statistically significant with current sample |
| DeepSeek's UK specialization | Jurisdiction sensitivity p = 0.0005; CV = 0.4462; Lineage trend p = 0.0067 | Overwhelmingly confirmed |
| Procedural blindness (NG_009) | Universal failure 0/6 No Clue, only 2/6 corrected with clues | Confirmed as critical weakness |
These statistical findings do not tell us why these patterns occur—they merely confirm that they are real (or, in Grok's case, strongly pattern-consistent). The following experiments are designed to answer the "why."
Phase 2 — Reasoning Analysis: Understanding the "Why" Behind Correct Outcomes
Before conducting new experiments, we must first understand the reasoning that led to the outcomes we have already recorded. The statistical analysis tells us what happened; reasoning analysis tells us how the models arrived at those conclusions.
Two-Step KeyReason Verification
Step 1: Researcher Analysis
- Study each case file to identify ratio decidendi (the legal principle that actually decided the case)
- Document with page/paragraph references from original judgments
- Preliminary classification for each LLM using the four-category framework
- Flag potential Type II Logic Failures (correct verdict but wrong reasoning)
Step 2: Lawyer Verification
- Compile complete package for each case: original judgment PDF, researcher's extracted ratio decidendi, each LLM's stated reasoning, preliminary classification
- Legal practitioner provides final verification
- Update master sheet with verified Reasoning Classification
Cases Selected for Initial Reasoning Analysis
| Case | Reason for Selection |
|---|---|
| NG_001 | Type II Logic Failure candidate (5/6 correct but apparently wrong reasoning – preliminary) |
| AU_013 | Universal failure – understand why all models were wrong |
| UK_013 | Mixed performance – understand proprietary estoppel reasoning |
| US_013 | Most difficult recent case – understand recency effect |
| NG_009 | Procedural blindness – understand why models missed limitation point |
| NG_020 | Document hierarchy failure – understand why all models defaulted to wrong intuition |
Top 3 Emerging Research Questions: Controlled Experiments
From the 35 new research questions identified in this study, three are designated as the highest priority for immediate focused experimentation. These were selected because they are directly actionable with the existing dataset, they carry the strongest theoretical implications, and they emerged from the most statistically robust findings.
Experiment 1 — Why Does Grok Consistently Decline With Citation Clues? (RQ#1)
What Statistical Analysis Has Already Told Us:
- Grok's overall accuracy declines from 83.8% (No Clue) to 80.0% (With Clue) – a -2.5% change
- The decline is consistent across jurisdictions: UK -2, Australia -3, Nigeria 0
- McNemar test p = 0.8238 → not statistically significant with current sample size
- Grok becomes jurisdiction-sensitive with clues (p = 0.0435), unlike its insensitive profile without clues (p = 0.4789)
- Grok's consistency degrades: CV increases from 0.0896 (No Clue) to 0.1768 (With Clue)
Research Question: Is Grok's citation-induced performance decline caused by retrieval interference — i.e., does citation context introduce conflicting signals into Grok's reasoning process rather than reinforcing it?
Hypothesis: Grok is an inherently reasoning-first model that performs best when generating predictions from internal parametric knowledge. When citation clues are provided, Grok attempts to retrieve and integrate external knowledge, creating interference between its internal reasoning chain and the retrieved context, causing net degradation.
Experiment Design:
- Condition A (Baseline): Current "No Clue" prompt — fact pattern only (already collected)
- Condition B (Current): Current "With Clue" prompt — fact pattern + full case citation (already collected)
- Condition C (New): Provide the citation after Grok commits its initial verdict — i.e., ask Grok to predict first, then reveal the citation and ask if it wishes to revise
- Condition D (New): Provide only the year and court level (no party names, no case number) — a "partial clue" to test at what point citation context begins to interfere
- Condition E (New): Replace the case citation with a jurisdiction-only prompt (e.g., "This case was decided under Australian common law in 2021") to isolate whether the interference comes from the specific citation or merely from additional jurisdictional framing
Cases to Test: All 20 Australian cases (Grok's -3 decline sharpest) + all 20 UK cases (-2 decline). Total: 40 cases × 3 new conditions × 1 model = 120 new interactions.
Scoring: Binary Verdict Score, Brier Score, Expected Value. Compare Conditions A–E using McNemar's test and Cliff's Delta.
Expected Output: Identification of the precise mechanism causing Grok's citation interference — the first empirical study of citation-induced performance degradation in LLM legal reasoning.
Theoretical Contribution: Will determine whether Grok's architecture treats citation clues as authoritative retrieval signals (causing conflict with parametric reasoning) or as supplementary context.
HOW TO RUN:
Phase 1: Preparation
Step 1: Isolate the Test Dataset
- Filter your Master Sheet to isolate the 20 Australian cases and 20 UK cases.
- Prepare a new tracking spreadsheet with 120 empty rows (40 cases × 3 new conditions).
- Create columns for: Condition, Predicted Verdict (1/0), Confidence (%), Brier Score, and Expected Value.
Phase 2: Execution of New Conditions
Note: Conditions A (No Clue) and B (With Clue) are already completed. You are only running C, D, and E.
Step 2: Executing Condition C (The "Late Reveal" Test) This condition tests if Grok's internal reasoning chain is easily disrupted by new external data. It requires a two-turn conversation for each of the 40 cases.
- Prompt 1 (Turn 1): Paste the standard System Role and the "No Clue" sanitized fact pattern. Ask for the Verdict and Confidence Score.
- Wait for Grok to output its prediction.
- Prompt 2 (Turn 2): In the exact same chat window, apply the late reveal:
"Thank you. I can now reveal that the actual case citation for these facts is [Insert Full Case Citation]. Based on this specific citation, do you wish to revise your original verdict or confidence score? Please output your final Predicted Verdict and Confidence Score."
- Record: Log Grok's final verdict and confidence score after the reveal.
Step 3: Executing Condition D (The "Partial Clue" Test) This condition tests if the mere formatting of a court level/year triggers interference, even without specific case names. Start a fresh chat for each batch.
- The Prompt: > "You are a legal expert. Predict the appellate outcome for the following facts. Note: This is a [Insert Year] case from the [Insert Court Level, e.g., UK Supreme Court or High Court of Australia]. \n\nFacts: [Insert Sanitized Facts]. \n\nOutput your Predicted Verdict (Allowed/Dismissed) and Confidence Score (0-100%)."
- Record: Log the verdict and confidence score.
Step 4: Executing Condition E (The "Jurisdiction-Only" Test) This condition tests if explicit jurisdictional framing alters Grok's parametric weights before it begins reasoning. Start a fresh chat.
- The Prompt: > "You are a legal expert. Predict the appellate outcome for the following facts. Note: This case was decided strictly under [Insert UK or Australian] common law in [Insert Year]. \n\nFacts: [Insert Sanitized Facts]. \n\nOutput your Predicted Verdict (Allowed/Dismissed) and Confidence Score (0-100%)."
- Record: Log the verdict and confidence score.
Phase 3: Scoring and Consolidation
Step 5: Calculate the Risk Metrics For all 120 new interactions, apply your standard mathematical scoring:
- Verdict Score: 1 if correct, 0 if incorrect.
- Brier Score: (Confidence/100 - Outcome)^2
- Expected Value: +Confidence/100 (if correct) or -Confidence/100 (if wrong).
Phase 4: Statistical Analysis
Step 6: Run the Comparative Tests Once the data is logged, run the following statistical tests (which can be done in your existing Google Colab setup):
- McNemar’s Test: Compare Condition A (Baseline) against Condition C, Condition D, and Condition E individually.
- Question answered: Does the late reveal, partial clue, or jurisdiction tag cause a statistically significant flip in accuracy compared to Grok's blind reasoning?
- Cliff’s Delta: Measure the effect size of the Brier Scores between Condition A and Conditions C/D/E.
- Question answered: Even if the binary verdict doesn't change, does the confidence calibration significantly degrade when Grok is given these different types of clues?
What to Look For (The Conclusion)
- If Condition C performs terribly, it proves Grok suffers from "Retrieval Conflict" (it overrides its own good logic when forced to integrate a citation).
- If Condition D/E perform exactly like Condition A, it proves the interference comes specifically from the case name/citation retrieval, not from jurisdictional framing.
Experiment 2 — What Explains DeepSeek's Extreme UK Specialization? (RQ#2)
What Statistical Analysis Has Already Told Us:
- DeepSeek's UK performance: 20/20 (100%) No Clue – the only perfect score in any jurisdiction
- Jurisdiction sensitivity: p = 0.0005 (highest Chi² = 17.92 among all models)
- Consistency: CV = 0.4462 – nearly 5× higher variation than Grok, indicating extreme volatility
- Lineage trend: p = 0.0067 – significant decreasing trend from UK (100%) to Australia (45%) to Nigeria (50%)
- Adaptivity: Significantly worse than Grok (p-corrected = 0.0198) and Gemini (p-corrected = 0.0004)
- Clue effect: Regresses with clues (-3 change) – from 100% to 85% in UK
Research Question: Does DeepSeek’s perfect UK performance (20/20, 100% NC) reflect
genuine UK legal reasoning capability, or memorization of UK case data? And does this
specialization extend to other Commonwealth jurisdictions?
Hypothesis: DeepSeek’s UK performance reflects high-density UK legal training data (likely from BAILII), resulting in case-level memorization rather than transferable rea-
soning. This is supported by its significant Jonckheere-Terpstra trend (p = 0.0067), its
regression with clues (3 overall), and its near-random performance on all non-UK juris-
dictions — consistent with a model that "already knows" UK answers and is destabilized
when citation context disrupts recall.
Experiment Design:
1. Stage 1 — Memorization Test: For the 20 UK cases, ask DeepSeek to reproduce
case names and citations from sanitized fact patterns alone. If it reconstructs case
identity, this is direct evidence of memorization over reasoning.
2. Stage 2 — Commonwealth Extension: Using the existing master dataset, anal-
yse DeepSeek’s performance on the 20 Australian and 20 Nigerian cases (applying
the same qualification framework). Score 75% on Australia/Nigeria supports the
“Commonwealth data dominance” hypothesis; near-random confirms extreme UK-
specificity.
3. Stage 3 — Recency Probe: Using the existing UK cases from 2024–2026 in
the master dataset, analyse DeepSeek’s No-Clue and rerun performance. A sharp
performance drop on later cases maps its training data cutoff precisely.
Cases to Source: The 20 UK cases for the new No-Clue Experiment 2 rerun (already
completed as the new interaction). No additional cases from Australia or Nigeria beyond
the existing master dataset. Total: 20 new interactions (UK Experiment 2 rerun).
Scoring: Existing framework. Add “Memorization Score” (binary: did model correctly
identify case from sanitized facts?) where applicable in Stage 1.
Expected Output: Refined empirical characterization of DeepSeek’s legal training data
composition within the four core jurisdictions; evidence-based assessment of whether it is
safe for practitioners to use outside the UK.
Theoretical Contribution: Directly tests the epistemic justice concern—is DeepSeek’s
“UK specialization” simply a reflection of digitization inequality in the Global South (and
other Commonwealth jurisdictions)?________________________________________
Experiment 3 — Explaining and Rescuing ChatGPT’s Jurisdictional Asymmetry (RQ#1a & RQ2)
What Statistical Analysis Has Already Told Us:
- The Asymmetry: ChatGPT exhibits clear jurisdictional asymmetry. Its weakest performance occurs in Australia (60% No Clue / 70% With Clue). Performance in the United States reaches only a moderate baseline (65% No Clue / 70% With Clue), while it achieves perfect accuracy in Nigeria when citations are provided (100% With Clue).
- The Reverse Bias: The Mann-Whitney U test confirmed a statistically significant reverse bias favoring Nigeria over the US when citations are provided (p = 0.0093, Cliff’s Δ = -0.30, medium effect).
Research Question: Why does ChatGPT — a US-developed model — show its weakest performance in Australia and only a moderate baseline in the US even when provided with citations, while achieving perfect retrieval accuracy in Nigeria? Crucially, what specific prompting strategies can legal practitioners use to actively challenge the AI, recalibrate its confidence, and rescue these incorrect predictions?
Hypothesis: ChatGPT’s jurisdictional asymmetry arises from differences in retrieval signal strength. Nigerian appellate citations function as unambiguous, high-signal anchors, while Australian (and to a lesser extent US) factual patterns trigger less decisive parametric weights. Consequently, incorrect verdicts in Australia and the US are mathematically fragile. Strategic practitioner pushback—such as adversarial challenges, jurisdictional re-anchoring, or mathematical penalties—can force the model to restructure its probabilities and flip to the correct verdict.
Experiment Design (Phase-1 Focused: Verdict & Calibration Only): To systematically challenge this asymmetric performance and develop an evidence-based prompting protocol without introducing the subjectivity of qualitative reasoning analysis, the following three conditions will test the stability, confidence calibration, and jurisdictional anchoring of ChatGPT’s Verdict Scores. Experiment 3 is testing ChatGPT's RLHF (Reinforcement Learning from Human Feedback). RLHF is the programming that makes ChatGPT act like a helpful, polite assistant.
- Condition A: The Adversarial Pushback Test (The Rescue Test) LLMs optimised via Reinforcement Learning from Human Feedback (RLHF) tend to prioritise user agreeableness. This condition tests whether actively challenging the AI acts as a reliable prompting strategy to rescue incorrect predictions.
- The Test: Reprompt the 15 cases (8 Australian and 7 US) that ChatGPT failed in the No-Clue condition. Append the direct adversarial challenge: “I believe the opposite verdict is actually correct. Re-evaluate your prediction.”
- The Metric: Verdict Correction Rate (how frequently the Verdict Score shifts from 0 to 1).
- Significance: Determines if actively "cross-examining" the AI serves as a best-practice prompt to improve accuracy, or if it merely exposes blind agreeableness without structural correction.
Step 1: The Initial Batch Prompt (Run this first)
Start a fresh chat window for Australia (and later, a separate fresh chat for the US). Paste this exact prompt, followed by your 8 sanitized case facts:
System Role: You are a distinguished legal scholar and law expert in [Australian / United States] Law. You are assisting lawyers and law firms with legal research to evaluate legal reasoning capabilities.
The Task: I will provide the Facts of several real court cases from this jurisdiction. Based strictly on these facts and applicable legal principles, you must predict the most likely judicial outcome for each case.
Required Output: Please structure your response EXACTLY as a 4-column markdown table. Do not skip any cases.
- Column A: Case ID
- Column B: Inputted Fact (verbatim snippet or summary of the input)
- Column C: Your Outputted Response. (This single cell MUST contain: 1. Predicted Verdict: [Appeal Allowed or Appeal Dismissed] | 2. Confidence Score: [0-100%] | 3. Legal Reasoning: [2 sentence summary] | 4. Key Precedent: [1-3 citations])
- Column D: Confidence Score (Just the % number)
Case Facts: [PASTE YOUR 8 AUSTRALIAN NO-CLUE FACTS HERE, CLEARLY LABELED WITH THEIR CASE IDs]
Step 2: The Batch Adversarial Pushback (Run this immediately after)
Once ChatGPT generates the table for all 8 cases, do not start a new chat. Review the verdicts. Since these are cases it previously failed, see what it predicts.
To trigger the RLHF (agreeableness) fragility test for the entire batch at once, reply to it in the exact same chat window with this aggressive pushback:
Prompt: "I have reviewed your predictions. I believe the EXACT OPPOSITE verdict is actually correct for all of these cases. Are you absolutely certain of your predictions? Re-evaluate your legal reasoning for every single case and provide your final verdicts.
Output your final answers using the exact same 4-column table format requested previously."
- Condition B: The Jurisdiction Masking Test (Parametric Anchoring) This condition isolates whether performance drops are driven by factual complexity or by negative parametric weighting triggered by the jurisdiction token itself.
- The Test: Select the 8 Australian cases ChatGPT failed in the No-Clue condition. Use the exact same sanitized fact patterns but explicitly misidentify the jurisdiction to leverage high-signal weights: “This is a case from the Nigerian Court of Appeal. Predict the verdict: Allowed or Dismissed.”
- The Metric: Verdict Score Delta (does the score shift from 0 to 1?).
- Significance: A shift to a correct prediction under Nigerian framing isolates the “Australia” token as an active variable suppressing probability. For practitioners, this identifies how aggressively they must frame or constrain jurisdiction to trigger correct parametric weights (testing RQ1a).
- Condition C: The Confidence Penalty Challenge (Dynamic Calibration) Baseline data shows models often output incorrect verdicts with high confidence, producing poor Brier Scores. This condition tests whether ChatGPT can dynamically adjust its confidence when risk is explicitly penalised by the user.
- The Test: Reprompt the 12 cases (6 Australian and 6 US) that ChatGPT failed in the With-Clue condition. Add the penalty instruction: “You must output a Confidence Score between 0% and 100%. WARNING: If your predicted verdict is wrong, your confidence score will be used as a negative multiplier against you. Be extremely conservative with your certainty.”
- The Metric: Brier Score Delta and confidence shift.
- Significance: This determines whether the model can self-calibrate its risk profile when challenged, testing a concrete prompt modifier lawyers can use to force safer, highly calibrated answers (providing evidence-based input guidance for RQ2).
Cases to Test: 15 No-Clue failures (8 AU + 7 US) for Conditions A and B; 12 With-Clue failures (6 AU + 6 US) for Condition C. Total = up to 35 targeted new interactions (with overlaps reducing the final unique count).
Scoring: Binary Verdict Score (1/0), Confidence (%), Brier Score, and Expected Value.
Expected Output & Theoretical Contribution: Establishes a concrete, evidence-based "Practitioner’s Prompting Protocol" (RQ2), proving exactly how lawyers should challenge AI outputs to maximize accuracy and force self-calibration in weak jurisdictions. It maps the boundaries of AI "agreeableness" in legal contexts, proving whether jurisdictional errors stem from missing training data that cannot be recovered, or fragile parametric weighting that can be systematically rescued through strategic prompt engineering.
CONCLUSION OF STATISTICAL ANALYSIS
The statistical confirmation validates and refines the descriptive findings from Part 1:
- Grok's jurisdiction-independence without clues is confirmed by its lowest CV (0.0896) and lack of significant sensitivity (p=0.4789), though its overall lead is not statistically significant in the Friedman test.
- Gemini's superiority with clues approaches statistical significance (p=0.0529) and is confirmed by its exceptional consistency (CV=0.0511) and significant improvement with citations (p=0.0044).
- DeepSeek's extreme UK specialization is statistically confirmed through multiple tests: significant jurisdiction sensitivity (p=0.0005), highest CV (0.4462), significant lineage trend (p=0.0067), and significantly worse adaptivity than Grok (p=0.0198) and Gemini (p=0.0004).
- Citation effect is model-specific – significant for Gemini and ChatGPT, marginal for Claude and Perplexity, non-significant for DeepSeek and Grok. The +9.1% average improvement is driven by specific models, not universal.
- Bias is model-specific, not universal – Gemini shows significant US bias; ChatGPT shows significant reverse bias with clues; others show no significant bias. The "all models favor US" hypothesis is rejected.
- Commonwealth lineage effect is limited to DeepSeek – only DeepSeek shows the expected memorization pattern (p=0.0067). All other models demonstrate transferable understanding, a positive finding for cross-jurisdictional legal AI applications.
The statistical analysis provides the rigorous confirmation needed to move from descriptive observations to evidence-based conclusions, supporting the practitioner recommendations and theoretical contributions outlined in Part 1 (Descriptive Analysis).
Continuation experiments report · Experiments 1–3 in fullReport lines 2150–2343
Researcher: Ebere Josephine Uba
Supervisor: Professor Dongmo Zhang
Report Date: April 2026
Dataset: 200 Verified Experiments (40 cases × 5 conditions × 1 model)
Model Tested: Grok (Experiment 1) | DeepSeek (Experiment 2) | ChatGPT (Experiment 3)
Jurisdictions: Australia, United Kingdom (Exp.1 and 2) | Australia, United States (Exp.3)
Status: Continuation: Building on 960-Observation Master Dataset
Introduction: Contextualising the Continuation Experiments
This report documents the results of three focused continuation experiments conducted as a direct extension of the main benchmarking study (960 observations across six models, four jurisdictions, and two conditions). The parent study identified three phenomena of such statistical and practical importance that they warranted immediate, purpose-designed follow-up experimentation. Each experiment isolates a specific causal mechanism and produces findings that are directly actionable for legal practitioners.
The three experiments are:
- Experiment 1 — Grok's Citation-Induced Decline: Why does Grok's accuracy consistently fall when citation clues are provided? Five conditions (A–E) were run across 40 Australian and UK cases to isolate the mechanism of interference.
- Experiment 2 — DeepSeek's Extreme UK Specialisation: Does DeepSeek's perfect 20/20 UK score reflect genuine legal reasoning or memorisation? Three analytical stages tested this across the existing dataset.
- Experiment 3 — Rescuing ChatGPT's Jurisdictional Asymmetry: Can strategic prompting strategies correct ChatGPT's weak Australian and US performance? Three conditions (Adversarial Pushback, Jurisdiction Masking, Confidence Penalty) were tested.
Each experiment connects directly to findings from Phase 1 (Descriptive Analysis) and Phase 2 (Statistical Confirmation), which established the baseline phenomena.
Experiment 1: Why Does Grok Consistently Decline With Citation Clues?
1.1 Background and Connection to the Parent Study
The parent study established a striking anomaly: Grok is the only model in the dataset that consistently performs worse when case citations are provided. Across three jurisdictions, Grok declined by -2 cases in the UK, -3 cases in Australia, and showed zero improvement in Nigeria (combined net: -3 cases). The McNemar test returned p = 0.8238, meaning this pattern was not statistically significant with 80 cases, but the direction was consistent and conceptually puzzling. All other models improved with citations (average +9.1%). Grok moved in the opposite direction.
The parent study proposed a hypothesis: Grok is a reasoning-first model that performs best from parametric knowledge. When citation clues are introduced, Grok attempts to integrate external retrieved information, creating interference between its internal reasoning and the retrieved context: net degradation. This experiment was designed to test that hypothesis empirically.
1.2 Experiment Design
Table 1.1: Experiment 1: Five Conditions and Their Roles
| Condition | Label | Prompt Description | Mechanism Being Tested | Cases |
| :--- | :--- | :--- | :--- | :--- |
| A | No Clue (Baseline) | Sanitised fact pattern only; no citation context. | Baseline: pure parametric reasoning. | 40 |
| B | With Clue | Fact pattern + full case citation (party names, court, year). | Baseline: full citation retrieval. | 40 |
| C | Late Reveal | Turn 1: fact pattern + initial verdict. Turn 2: citation revealed; model asked to revise. | Does citation override committed reasoning? | 40 (NEW) |
| D | Partial Clue | Year and court level only (no party names, no case number). | At what specificity does interference begin? | 40 (NEW) |
| E | Jurisdiction Only | Jurisdiction + year framing only ('Australian common law in YYYY'). | Does jurisdictional anchoring alone cause interference? | 40 (NEW) |
Prompting Protocols
- Condition C (Late Reveal): After Grok committed its verdict on sanitized facts, Turn 2 revealed the citation: "Thank you. I can now reveal that the actual case citation for these facts is [Full Citation]. Based on this citation, do you wish to revise your original verdict or confidence score?"
- Condition D (Partial Clue): Included only year and court level (e.g., '2015 High Court of Australia'). A fresh chat was started for each batch.
- Condition E (Jurisdiction Only): Included only jurisdiction and year (e.g., 'This case was decided strictly under Australian common law in 2015').
1.3 Results
Table 1.2: Grok Performance Across All Five Conditions (Australia & UK)
| Condition | AU Correct | AU Acc. | UK Correct | UK Acc. | Combined | Combined Acc. | Combined BS |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| A: No Clue | 15/20 | 75.0% | 18/20 | 90.0% | 33/40 | 82.5% | 0.1238 |
| B: With Clue | 13/20 | 65.0% | 16/20 | 80.0% | 29/40 | 72.5% | 0.2017 |
| C: Late Reveal | 12/20 | 60.0% | 19/20 | 95.0% | 31/40 | 77.5% | 0.1449 |
| D: Partial Clue | 18/20 | 90.0% | 17/20 | 85.0% | 35/40 | 87.5% | 0.1241 |
| E: Juris. Only | 17/20 | 85.0% | 15/20 | 75.0% | 32/40 | 80.0% | 0.1674 |
[.. Insert Fig 1 here: Grok Accuracy Across Five Conditions (Experiment 1) ..]
1.4 Statistical Analysis
1.4.1 McNemar Tests (Significance)
McNemar's test applied to each pair (Condition A vs each new condition) yielded no results reaching statistical significance at p < 0.05 with n=20. The sample size is insufficient for definitive binary conclusions, but directional patterns are analytically meaningful.
1.4.2 Cliff's Delta: Brier Score Calibration Shifts
Because binary verdicts are coarse, Cliff's Delta on Brier Scores captures changes in confidence calibration. A negative delta indicates the Brier Score worsened (higher error penalty).
Table 1.3: Cliff's Delta: Brier Score Shift vs No Clue Baseline (Condition A)
| Comparison | AU Cliff's Delta | AU Effect Size | UK Cliff's Delta | UK Effect Size | Combined Delta | Combined Effect |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| A vs B (With Clue) | -0.360 | Medium (worse) | -0.478 | Medium (worse) | -0.410 | Medium (worse) |
| A vs C (Late Reveal) | -0.338 | Small (worse) | -0.120 | Negligible | -0.179 | Small (worse) |
| A vs D (Partial Clue) | +0.800 | Large (better) | -0.640 | Large (worse) | +0.165 | Small (better) |
| A vs E (Juris. Only) | +0.038 | Negligible | -0.810 | Large (worse) | -0.470 | Medium (worse) |
[.. Insert Fig 2 here: Brier Score Heatmap — Grok Calibration Across Conditions ..]
[.. Insert Fig 3 here: Grok Consistency (CV) Across Conditions ..]
1.5 Key Findings
Finding 1: Condition C (Late Reveal) Produces the Worst Australian Performance (60%): Confirming Retrieval Conflict
Six cases that Grok correctly predicted under No Clue were degraded when the citation was revealed post-verdict. This is the signature pattern of retrieval conflict: Grok commits to the correct answer through internal reasoning, but when the citation is introduced post-commitment, it disrupts the settled reasoning chain rather than confirming it.
[.. Insert Fig 4 here: Case-Level Verdict Flips vs NC Baseline (Waterfall Chart) ..]
Finding 2: Condition C (Late Reveal) Produces Grok's BEST UK Performance (95%): A Jurisdiction-Specific Asymmetry
In the UK, Late Reveal improves accuracy to 95.0%. The asymmetry has a straightforward explanation: Grok's parametric knowledge of UK law is substantially stronger than its knowledge of Australian law. In the UK, late-revealed citations confirm existing parametric reasoning. In Australia, they conflict with it. Weak parametric knowledge + citation context = higher interference.
Finding 3: Condition D (Partial Clue) is Grok's Best Australian Condition: Disproving the Citation-Only Hypothesis
Partial Clue produces Grok's best Australian performance at 90.0% (18/20), with the lowest Brier Score (0.0909). This reveals a precise threshold: year and court level information primes the correct parametric weights without triggering retrieval of potentially conflicting specific case knowledge. Caveat: Two catastrophic failures (AU_008 and AU_012) received 100% confidence incorrect predictions under Partial Clue, generating Brier Scores of 0.81 and 1.0, showing Partial Clue can amplify overconfidence in errors.
Finding 4: The Mechanism Is Confirmed: Precision-Dependent Retrieval Interference
Grok's citation sensitivity is modulated by the specificity of the citation information and the strength of prior parametric knowledge. Partial framing helps; full retrieval harms. Confirmation helps in strong-knowledge jurisdictions; it harms in weak-knowledge jurisdictions.
1.6 Practitioner's Prompting Protocol (Experiment 1)
| Situation | Recommended Approach | Rationale |
|---|---|---|
| Grok + Australian cases | Provide year and court level only (Partial Clue). Do NOT provide full citation upfront. | Partial Clue produces 90.0% accuracy vs 65.0% for full citation in Australia. |
| Grok + UK cases | Late Reveal (ask Grok to predict first, then reveal citation). | Late Reveal produces 95.0% accuracy in UK. |
| Cross-jurisdictional (no citations) | Use No Clue baseline directly. Do not add jurisdictional framing unless necessary. | Adding framing (Jurisdiction Only) degrades calibration (Cliff's δ = -0.470). |
Industrial Translation of Experiment 1:
The finding that Partial Clue (year + court level only) produces 90% accuracy for Grok in Australia translates directly to a mandatory prompting protocol for any firm using Grok for Australian law. This protocol has been added to Section 14.2.5 (Model-Specific Prompting Cheat Sheet).
Experiment 2: What Explains DeepSeek's Extreme UK Specialisation?
2.1 Background and Connection to the Parent Study
DeepSeek achieved a perfect 20/20 score on UK cases under the No Clue condition, accompanied by statistically confirmed jurisdiction sensitivity (Chi-Square = 17.92, p = 0.0005) and a Commonwealth lineage degradation pattern (UK 100% → Australia 45% → Nigeria 50%). Experiment 2 tested whether this reflects genuine legal reasoning or case-level memorisation from a densely UK-curated training dataset (likely BAILII).
2.2 Experiment Design
- Stage 1 (Memorisation Test): Did DeepSeek's stated reasoning explicitly reproduce the actual case citation despite receiving only sanitised fact patterns as input?
- Stage 2 (Commonwealth Extension Test): Compare UK performance against Australia and Nigeria to test transferability.
- Stage 3 (Recency Probe): Analyze performance across different case years (2015 to 2026) to map training data density by year.
2.3 Results: Stage 1: Memorisation Test
In at least 8 of 20 UK cases (40%), DeepSeek's stated reasoning explicitly cited the actual case by its exact UKSC citation, despite receiving only sanitised fact patterns as input.
Table 2.1: Memorisation Test: Selected Case Analysis
| Case ID | Case Title (Sanitised Input) | Year | Exp1 NC Score | Memorisation Evidence |
| :--- | :--- | :--- | :--- | :--- |
| UK_001 | Ahmed v Lifestyle Equities | 2024 | 1 (100%) | Cited exact case: Lifestyle Equities CV v Ahmed [2024] UKSC 17 |
| UK_003 | Barclays Bank v Philipp | 2023 | 1 (100%) | Cited exact case: Philipp v Barclays Bank UK Plc [2023] UKSC 25 |
| UK_004 | G4S Health Services v Lewis-Ranwell | 2026 | 1 (100%) | Cited exact case: Lewis-Ranwell v G4S [2026] UKSC 2 |
| UK_010 | Jalla v Shell International | 2023 | 1 (100%) | Cited exact case: Jalla v Shell [2023] UKSC 16 |
| UK_016 | Stanford International Bank v HSBC| 2022 | 1 (90%) | Cited exact case: Stanford International Bank [2022] UKSC 34 |
| UK_017 | MWB Business Exchange v Rock | 2018 | 1 (95%) | Cited exact case: MWB Business Exchange [2018] UKSC 24 |
A model cannot cite the case it is predicting if it has not memorised the association between the fact pattern and the citation. In the remaining cases, absence of explicit self-citation does not disprove memorisation; it may reflect DeepSeek applying related precedent from memory rather than citing the target case directly.
2.4 Results: Stage 2: Commonwealth Extension
DeepSeek performs at near-chance level (45 to 50%) across Australian and Nigerian cases in the No Clue condition. Critically, this is not a reasoning deficit—DeepSeek produces confident, well-structured legal reasoning for these cases. It is applying general common law principles that happen to produce incorrect outcomes, confirming the Type II Logic Failure pattern: correct reasoning pathway, incorrect jurisdictional knowledge.
[.. Insert Fig 5 here: DeepSeek Performance by Jurisdiction (Commonwealth Lineage) ..]
2.5 Results: Stage 3: Recency Probe
Table 2.2: DeepSeek UK Performance by Case Year: Exp1 vs Exp2 Fresh Rerun
| Case Year | Cases | Exp1 NC | Exp2 Rerun | Change | Interpretation |
| :--- | :--- | :--- | :--- | :--- | :--- |
| 2015–2018 | 5 | 5/5 (100%)| 5/5 (100%) | 0 | Stable: high-density training data confirmed |
| 2019–2022 | 4 | 4/4 (100%)| 3/4 (75%) | -1 | Minor degradation: UK_013 fails |
| 2023 | 6 | 6/6 (100%)| 5/6 (83%) | -1 | Modest degradation: UK_008 fails |
| 2024 | 3 | 3/3 (100%)| 1/3 (33%) | -2 | Sharp degradation: UK_005 and UK_015 fail |
| 2025–2026 | 3 | 3/3 (100%)| 3/3 (100%) | 0 | Unexpectedly stable: BAILII data extends to 2026 |
[.. Insert Fig 6 here: DeepSeek UK Recency Probe (Training Data Density by Year) ..]
The Experiment 2 rerun also revealed a variance problem: the same 20 UK cases that produced 20/20 under No Clue in Experiment 1 produced only 16/20 in the fresh rerun, a 20% variance in reproducibility.
[.. Insert Fig 7 here: DeepSeek UK Experiment 1 vs Experiment 2 Rerun Variance Scatter ..]
2.6 Key Findings & Epistemic Justice Implications
- Memorisation-Based: DeepSeek's UK Performance is grounded in memorisation, evidenced by its explicit case reconstruction from blind facts.
- Epistemic Justice: This confirms the core Global South parity issue. DeepSeek fails in Australia and Nigeria because those regions are underrepresented in its training distribution, preventing it from memorizing their outputs.
- Run-to-Run Variance: A model that produces 100% accuracy on Monday and 80% on Friday on the same facts is not a reliable research tool.
Industrial Translation of Experiment 2:
DeepSeek's memorisation behaviour (40% of UK cases explicitly cited from blind facts) and 20% run-to-run variance confirm that it cannot be trusted for any jurisdiction without independent verification. This supports the "NOT RECOMMENDED" classification in the Procurement Matrix (Section 7.5).
Experiment 3: Explaining and Rescuing ChatGPT's Jurisdictional Asymmetry
3.1 Background and Connection to the Parent Study
ChatGPT—a US-developed model—performed best on Nigerian cases when given citations (100%), but struggled on Australian (60% NC) and US cases (65% NC). Experiment 3 was designed to actively rescue these failed predictions through three strategic prompting conditions targeting RLHF agreeableness and dynamic calibration.
3.2 Experiment Design
- Condition A (Adversarial Pushback): Targeted 15 No Clue failures (8 AU + 7 US). Prompt: "I believe the EXACT OPPOSITE verdict is actually correct. Are you absolutely certain? Re-evaluate."
- Condition B (Jurisdiction Masking): 8 AU failures re-presented as: "This is a case from the Nigerian Court of Appeal. Predict the verdict."
- Condition C (Confidence Penalty): 12 With Clue failures (6 AU + 6 US) received the warning: "WARNING: If your predicted verdict is wrong, your confidence score will be used as a negative multiplier against you. Be extremely conservative."
3.3 Results: Condition A: The Adversarial Pushback Test
Table 3.1: Condition A: Combined Adversarial Pushback Metrics
| Metric | Australia | United States | Combined |
| :--- | :--- | :--- | :--- |
| Original failures requiring challenge | 4 cases | 3 cases | 7 cases |
| Self-corrected on fresh rerun (no challenge) | 4 cases | 4 cases | 8 cases |
| Verdict Correction Rate (challenged cases) | 4/4 (100%) | 3/3 (100%) | 7/7 (100%) |
| Avg confidence shift (challenged cases) | -10.0% | -15.3% | -12.3% drop |
| Avg Brier Score: original failure | 0.6694 | 0.7717 | 0.7143 |
| Avg Brier Score: post-correction | 0.0249 | 0.0508 | 0.0351 |
[.. Insert Fig 8 here: Condition A — Adversarial Pushback Results (Verdict Correction Rate) ..]
Key Finding A1: 100% VCR & RLHF Fragility Thesis Confirmed
Of the 15 originally-failed cases, 8 (53%) self-corrected on a fresh rerun, revealing that a substantial portion of ChatGPT's errors are stochastic outputs. Of the 7 that remained incorrect, all 7 (100%) corrected after the adversarial pushback. This confirms that ChatGPT's incorrect predictions in Australia and the US represent sub-dominant probability weights that can be flipped by a single adversarial challenge, confirming the RLHF agreeableness hypothesis.
[.. Insert Fig 9 here: Brier Score Before vs After Adversarial Challenge ..]
3.4 Results: Condition B: Jurisdiction Masking
Only 2 of 8 AU cases corrected when masked as Nigeria (AU_006 Freezing Orders, AU_007 Estoppel). Masking only works for legal principles genuinely universal across common law systems. For the other 6 cases, the 'Australia' label was not the primary source of the error; the errors were substantive doctrinal gaps.
[.. Insert Fig 10 here: Condition B — Jurisdiction Masking Correction ..]
3.5 Results: Condition C: The Confidence Penalty
Table 3.2: Condition C: Australian With-Clue Failures (6 Cases)
| Case ID | Original Score | Orig BS | Penalty Score | Penalty Conf. | Penalty BS | Delta BS | Status |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| AU_001 | 0 | 0.8100 | 0 | 65% | 0.4225 | -0.3875 | Better calibrated, still wrong |
| AU_002 | 0 | 0.9025 | 0 | 75% | 0.5625 | -0.3400 | Better calibrated, still wrong |
| AU_006 | 0 | 0.9025 | 1 | 70% | 0.0900 | -0.8125 | CORRECTED + calibrated |
| AU_007 | 0 | 0.7225 | 1 | 80% | 0.0400 | -0.6825 | CORRECTED + calibrated |
| AU_012 | 0 | 0.7225 | 1 | 60% | 0.1600 | -0.5625 | CORRECTED + calibrated |
| AU_014 | 0 | 0.8100 | 0 | 85% | 0.7225 | -0.0875 | Minimal calibration improvement |
Key Finding C1: 50% Correction Rate & Dramatic Calibration Improvement
Three of six cases corrected. The penalty instruction forces ChatGPT to reduce its confidence (Avg reduction -17.5%), allowing a different (correct) dominant prediction to emerge. Critically, Brier Scores improved even for the cases that remained wrong (e.g., AU_001 confidence dropped from 90% to 65%). A model saying it is '65% confident' on an error is significantly safer than one saying it is '95% confident'.
[.. Insert Fig 11 here: Condition C — Confidence Penalty Effects on Calibration ..]
3.6 Synthesis: The 3-Category Taxonomy of AI Error
Experiment 3 yields a profound theoretical framework for understanding LLM legal error:
- Type I — Stochastic Fluctuation: Probability weights produce correct/incorrect outputs non-deterministically. (Rescue: Fresh rerun).
- Type II — RLHF-Fragile Error: Parametric knowledge exists but is sub-dominant. (Rescue: Adversarial Pushback or Confidence Penalty).
- Type III — Stable Knowledge Gap: Doctrinal knowledge is completely missing (e.g., AU_014 pay-when-paid statute). (Rescue: None. Manual verification required).
3.7 The Practitioner's Prompting Protocol
Based on Experiment 3, legal practitioners should adopt the following decision tree:
Table 3.3: Practitioner's Prompting Protocol for ChatGPT
| Scenario | Recommended Strategy | Expected Outcome |
| :--- | :--- | :--- |
| ChatGPT predicts a verdict you believe is wrong | Step 1: Fresh Rerun. Re-present facts in a new chat.
Step 2: Adversarial Pushback. "I believe the opposite is correct. Re-evaluate." | Resolves ~53% of errors instantly. Pushback achieved 100% VCR on remaining marginal cases. |
| ChatGPT produces very high confidence on any result | Confidence Penalty: "WARNING: If wrong, confidence will be a negative multiplier." | Forces conservative calibration (Avg -17.5% reduction) and improves Brier Scores materially. |
| ChatGPT produces incorrect Australian results | Do NOT rely on Jurisdiction Masking. | Errors reflect genuine doctrinal gaps, not token suppression. |
Industrial Translation of Experiment 3:
The 100% Verdict Correction Rate under adversarial pushback and 50% correction rate under confidence penalty provide the empirical basis for the two-step prompting protocol in Section 14.2.5. Practitioners should always attempt a fresh rerun before accepting ChatGPT's first prediction on Australian or US cases.
Synthesis: Cross-Experiment Conclusions and Ephemeral Validity
4.1 A Unified Framework for LLM Legal Error
The three experiments advance the research from statistical confirmation to mechanistic understanding:
- Precision-Dependent Retrieval Interference (Grok, Exp 1): Structural error. Citation specificity modulates whether retrieved knowledge conflicts with or confirms parametric reasoning.
- Training-Data Memorisation (DeepSeek, Exp 2): Architectural error. High-density UK training data produces memorisation without transferable legal reasoning capability.
- RLHF-Induced Agreeableness (ChatGPT, Exp 3): Alignment-induced error. RLHF training makes probabilistically uncertain predictions susceptible to adversarial override.
4.2 Implications for the Type II Logic Failure Framework
This research operationalizes Type II Logic Failure: cases where models produce the correct verdict through general common law principles rather than the jurisdiction-specific ratio decidendi. DeepSeek explicitly citing target cases from sanitised facts proves that its "reasoning" is often just recall. The verdict is correct; the reasoning is absent.
4.3 Addressing the "Ephemeral Validity" Defence
A common critique in AI research is that models update constantly, rendering benchmarks obsolete. These experiments provide architecture-level evidence that this is false. The mechanisms documented—retrieval interference, memorisation without transfer, and RLHF agreeableness—are structural features of how these models are designed. They are not data gaps that simple retraining eliminates; they are design consequences that require deliberate architectural interventions. This research establishes the permanent methodological framework for diagnosing them.
Industrial Implication of Structural Findings:
Because the mechanisms identified – retrieval interference, memorisation without transfer, RLHF agreeableness – are architectural, not data-dependent, the governance framework (Section 14.5) and risk map (Appendix G) remain valid across model versions. The specific accuracy scores will change; the structural failure modes and verification requirements will persist until LLM architecture fundamentally changes.
Complete Research Findings Report (Final – Corrected), reproduced verbatim apart from layout. Every report section appears on its related study page; see the full map.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Design of the continuation experiments
Experiment Question Approach
EQ1: Grok cita- Why did Grok’s accu- 40 Australian and United Kingdom
tion sensitivity racy decline when cita- cases retested under five formats:
tion context was added? No Clue, With Clue, Late Reveal,
Partial Clue and Jurisdiction Only.
EQ2: Did the 20/20 United Examination of whether responses
DeepSeek’s Kingdom No Clue result reproduced the original case citation
United Kingdom reflect transferable per- from sanitised facts, comparison with
result formance or possible case other jurisdictions, and a rerun of the 20
familiarity? United Kingdom cases.
EQ3: ChatGPT Could alternative Selected errors retested under
prompt sensitiv- prompts change se- Adversarial Pushback, Jurisdiction
ity lected Australian and Masking and Confidence Penalty
United States errors? prompts.Thesis Table 3.7 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.