Key takeaways
- 76 numbered tables are reproduced on their evidence pages.
- Four numbered figures are represented by native charts or diagrams.
- Table transcriptions are source-reported and retain thesis layout and notes.
What was tested & why
All tables & figures
Select a numbered item below to jump to its evidence page. Numbers and captions refer to the submitted thesis, not independent re-analysis.
Source: Thesis Chapters 3–12 · source-reported unless otherwise noted.
Evidence index
76 tables & four figures
Table 3.1Phases of the CrossLaw project and their statusTable 3.2Primary case sources by jurisdictionTable 3.3Seven-criterion CrossLaw case qualification frameworkTable 3.4Case sanitisation protocolTable 3.5Prompt architecture of the primary CrossLaw experimentTable 3.6Inferential and supplementary methods used in the primary benchmarkTable 3.7Design of the continuation experimentsTable 4.1Verdict accuracy by model and jurisdiction under No Clue (correct out of 20 per jurisdiction)Table 4.2Verdict accuracy by model and jurisdiction under With Clue (cor- rect out of 20 per jurisdiction)Table 4.3Commonwealth-lineage comparison under No Clue (Jonckheere– Terpstra test of the ordering UK, Australia, Nigeria)Table 4.4Citation-context effect by model (paired change in correct predictions out of 80)Table 4.5Expected Calibration Error by model and prompting conditionTable 4.6Brier score by model and prompting conditionTable 4.7Silent Failure summary across the primary benchmarkTable 4.8Silent Failure rate by jurisdiction, aggregated across models and conditions (240 predictions each)Table 4.9Selected high-confidence incorrect predictionsTable 4.10Cases predicted incorrectly by all six models under No ClueTable 4.11Average No Clue accuracy across models by year of decision (%)Table 4.12Principal statistically significant results of the primary benchmarkTable 5.1First-stage experimental conditions (simultaneous loading)Table 5.2Second-stage experimental conditions (materials preloaded before the prediction task)Table 5.3Third-stage experimental conditions (shorter relational summaries preloaded, six models)Table 5.4Fourth-stage experimental conditions (stored summaries from five tools, ChatGPT as the fixed predictor)Table 5.5General Nigerian legal materials used in the added conditionsTable 6.1Model accuracy across the five first-stage experimental conditionsTable 6.2Model accuracy with fuller general Nigerian legal materials under No ClueTable 6.3Model accuracy with fuller general Nigerian legal materials under With ClueTable 6.4GLMRS No Clue model-level resultsTable 6.5Pooled accuracy across general-material No Clue presentationsTable 6.6Model accuracy across the two general-material No Clue presenta- tionsTable 6.7Pooled comparison of the original and added first-stage conditionsTable 7.1Model accuracy across the seven preloaded conditions: accuracy (number correct out of 20)Table 7.2Pooled results for the seven preloaded conditions (120 observations each)Table 7.3Model-level results for the four GLMP conditions (fuller materials preloaded, run 1 = a, rerun = b)Table 7.4Model-level results for the three GLMRS conditions (relational summaries preloaded, No Clue)Table 7.5Model accuracy by relational-summary length (No Clue), with preloaded fuller materials for referenceTable 7.6Pooled accuracy: simultaneous loading (first stage) versus preload- ing (second stage)Table 7.7Preloaded No Clue comparison by representation and summary length (differences in percentage points)Table 7.8Effect of adding the case-specific clue under preloading: model- level accuracy and number of verdicts that changed (out of 20)Table 7.9Run-to-run stability: verdict agreement between run a and the identical rerun b (out of 20 cases)Table 7.10Mean stated confidence by model and preloaded six-model condi- tionTable 7.11Sensitivity of the GLMP results (pooled accuracy, and clue effect in percentage points)Table 7.12Exploratory paired comparisons on the same 120 model–case pairs (discordant pairs and unadjusted exact McNemar p)Table 8.1Model-level results for the two third-stage GLMRS conditions (shorter relational summaries preloaded, No Clue)Table 8.2Pooled results for the two third-stage conditions (120 observations each)Table 8.3Class-level behaviour in the third-stage conditions (14 Dismissed and 6 Allowed cases per model)Table 8.4Pooled summary-length curve (No Clue, 120 observations per point)Table 8.5Summary-length curve by model: accuracy (%) over 20 casesTable 8.6GLMRS-A (50–100 words) against GLMRS-B (100–150 words): verdict agreement and accuracy change by modelTable 9.1Stored summaries by tool: summaries produced and word-range compliance (summary text only, excluding the title)Table 9.2ChatGPT prediction results by summarisation tool at 100–150 words per source (run a)Table 9.3ChatGPT prediction results by summarisation tool at 150–300 words per source (run a)Table 9.4Case-level results by summarisation tool (run a): number of tools (out of 5) whose summaries led to the correct verdict, and the tools whose summaries led to an errorTable 9.5100–150 against 150–300 words per source, by summarisation tool (run a)Table 9.6Run a against run b for the ChatGPT- and Claude-summary con- ditionsTable 10.1Pooled accuracy across all fourteen six-model conditions (120 observations each, majority-class reference = 84/120 = 70.0%)Table 10.2Model-level totals across the first-stage (five conditions), second- stage (seven conditions) and third-stage (two conditions) experimentsTable 10.3Pooled comparison of all ten six-model No Clue conditions (120 observations each)Table 10.4No Clue conditions grouped by input type (exploratory pooling of related experiments)Table 10.5Model accuracy (%) across all ten six-model No Clue conditionsTable 10.6Change from each model’s original No Clue accuracy (percentage points)Table 10.7Reliability indicators for the No Clue conditions (where reported)Table 10.8Pooled comparison of all four With Clue conditions (120 obser- vations each)Table 10.9Model accuracy (%) across all four With Clue conditionsTable 10.10Effect of the case-specific clue by model (With Clue minus paired No Clue, percentage points)Table 10.11Reliability indicators for the With Clue conditions (where re- ported)Table 10.12Class-level behaviour in the nine preloaded six-model conditions (Ground Truth: 14 Dismissed and 6 Allowed cases per model)Table 10.13Case-level results: number of models (out of 6) predicting the Ground Truth in each preloaded six-model conditionTable 10.14Mapping of the research questions to experiments, results, dis- cussion and conclusionTable 12.1Mapping between the labels used in this thesis and the original labelsTable 12.2Fourth stage, run a, 100–150 words per source: ChatGPT verdicts by summarisation toolTable 12.3Fourth stage, run a, 150–300 words per source: ChatGPT verdicts by summarisation toolTable 12.4Fourth stage, run b (reruns): ChatGPT verdicts with ChatGPT and Claude summariesTable 12.5Word count of each stored summary (summary text only, exclud- ing the title)Table 12.6Benchmark cases and Ground TruthTable 12.7Project resourcesFigure 3.1Case qualification funnel: the three-step selection processFigure 4.1Overall accuracy of each model under No Clue and With ClueFigure 8.1Accuracy against summary length for each model and pooledFigure 9.1ChatGPT’s accuracy by summarisation tool and summary length