CrossLaw/Putting it together/Answers to the 11 research questions
17 / 26 · Putting it togetherThesis §10.9; Table 10.14

Each question has an evidence path.

Answers to the 11 research questions

Key takeaways

  • RQ1–RQ2: model, jurisdiction, citation and confidence.
  • RQ3–RQ10: representation, order, length, baseline and stability.
  • RQ11: summarisation-tool differences with a fixed predictor.

What was tested & why

Answers to the 11 research questions

The thesis answers RQ1–RQ11 in §10.9 and maps each question to experiments and results in Table 10.14. Use the linked evidence pages below to inspect each answer.

Source: Thesis §10.9; Table 10.14 · source-reported unless otherwise noted.

Thesis §10.9 · source-reported

Answers and evidence

RQ1

How do verdict accuracy, confidence reliability and Silent Failures vary across six commercial LLMs when they predict the outcomes of final appellate cases from Nigeria, Australia, the United Kingdom and the United States?

They varied substantially. Grok had the highest aggregate No Clue accuracy (83.8%) and Gemini the highest With Clue accuracy (93.8%), so no model led in both conditions. Models varied across jurisdictions, most clearly DeepSeek (20/20 in the United Kingdom against 9 or 10 of 20 elsewhere under No Clue), and only DeepSeek showed a significant ordered decline from the United Kingdom to Australia to Nigeria. DeepSeek also had the highest ECE and Brier score in both conditions. Of the 960 predictions, 222 (23.1%) were Silent Failures, and the rate was highest for Australia (30.8%), then Nigeria (23.8%), the United States (22.5%) and the United Kingdom (15.4%). Performance therefore did not follow a simple legal-resource hierarchy.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ2

Does adding citation-based case context to the prompt change verdict accuracy and Silent Failures compared with sanitised facts alone, and does the effect vary across models and jurisdictions?

Pooled accuracy rose from 72.1% to 81.3%, and Silent Failures fell from 134 to 88. Five models improved and Grok declined, and the paired change was statistically significant for Gemini and ChatGPT only. The pooled gain was largest for Nigeria (+16.7 points) and smallest for the United States (+3.3 points). The Grok continuation experiment suggested that the presentation of the context also mattered. The effect was therefore model-dependent and did not establish why models changed their answers.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ3

Does adding broad Nigerian legal materials change outcome prediction relative to the original No Clue and With Clue conditions?

Under No Clue, the fuller materials were associated with higher pooled accuracy than facts alone: 72.5% under simultaneous loading and 75.0% and 74.2% when preloaded, against 67.5%. Under With Clue, no material-based condition reached the original With Clue baseline (75.8% to 80.0% against 84.2%). The materials therefore narrowed the gap between the two baselines but did not substitute for the case-specific clue.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ4

When the target case remains under No Clue, does presenting the same general legal materials as connected relational summaries produce different results from presenting the fuller materials?

Under simultaneous loading they gave the same pooled accuracy (72.5% each). Under preloading, model-constructed summaries at five lengths gave 64.2% to 67.5%, below the preloaded fuller materials (74.6%). The answer therefore depended on the loading protocol and on how the summaries were produced.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ5

Do any observed effects remain consistent across the six models?

No. In both stages the same condition was associated with higher accuracy for some models and lower accuracy for others. Grok was below its own factsonly baseline in all nine material-based No Clue conditions and Perplexity in seven of them, whereas ChatGPT, Gemini, Claude and DeepSeek were above their baselines in most conditions.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ6

Does preloading the materials before the prediction task change accuracy relative to supplying them simultaneously with the task?

The preloaded fuller-material runs were 1.7 to 4.2 points above the corresponding simultaneous-loading results. The direction was consistent, but the size was within the observed run-to-run variation, and the comparison crosses stages that were run at different times.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ7

Are preloaded results stable when the identical experiment is rerun?

At pooled level, largely yes (75.0% and 74.2% under No Clue, 80.0% and 78.3% under With Clue). At verdict level, agreement was about 75%, with model-level swings of up to 55 points. The fixed-predictor reruns of the fourth stage agreed on 18 or 19 of 20 verdicts. Single-run evaluation may therefore provide an incomplete picture of model behaviour.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ8

Does the permitted length of the relational summaries affect accuracy and reliability? The second stage tested 150–300, 300–500 and 500–750 words per source, and the third stage extended the range downwards to 50–100 and 100–150 words.

No reliable pooled effect was observed. Pooled accuracy was 65.8%, 65.8%, 67.5%, 65.8% and 64.2% from 50–100 to 500–750 words, and no pairwise difference was distinguishable from chance. Within the tested configurations, pooled accuracy did not increase monotonically with permitted summary length. The confidence–accuracy gap was close to zero at the two shortest lengths and larger at the longer lengths.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ9

How do all conditions compare with the majority-class baseline, and how do results differ across case outcomes and individual cases?

Only the conditions that included the clue and the first preloaded No Clue run exceeded the 70.0% always-Dismissed reference by five points or more. Across the nine preloaded six-model conditions, accuracy was 77.4% on Dismissed cases and 55.2% on Allowed cases. Difficulty was concentrated in a few cases, notably NG 009, NG 006, NG 017, NG 019, NG 020 and NG 011.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ10

Does each model respond to summary length in the same way, or does each model have its own length curve between no added context and the full materials?

No. The highest observed accuracy occurred at 50–100 words for DeepSeek, at 100–150 words for ChatGPT, at 150–300 words for Gemini and at 150– 300 words or longer for Claude. Grok and Perplexity were below their facts-only baselines at every summary length.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

RQ11

At a fixed summary length, does the tool used to produce the relational summaries change the accuracy, reliability and run-to-run consistency of a fixed prediction model?

With ChatGPT as the predictor, accuracy ranged from 75.0% to 90.0% across the five tools at each length, with the highest observed accuracy for ChatGPT’s and Claude’s summaries. The differences arose on five cases and are not statistically separable on 20 cases. Every tool gave equal or higher accuracy at 100–150 words than at 150–300 words. ChatGPT was under-confident in every tool condition, and the reruns of the ChatGPT- and Claude-summary conditions agreed on 18 or 19 of 20 verdicts.

Thesis §10.9 · Table 10.14 · source-reported

Read the relevant result

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 10.14 · source-reported

Mapping of the research questions to experiments, results, dis- cussion and conclusion

RQ   Topic                     Part         and Results                      Discussion and
                               stage(s)                                      conclusion

1    Model, jurisdiction and Part I                 Ch. 4                    Sec. 4.10, 10.6, Ch. 11
     confidence                                                              (Findings 1 and 2)
2    Citation context           Part I, continua-   Ch. 4                    Sec. 4.10, 10.6, Ch. 11
                                tion experiments                             (Findings 1 and 2)
3    Broad materials versus Part II, first and      Ch. 6, Sec. 10.2, 10.3   Sec. 10.7, Ch. 11 (Finding 3)
     baselines                  second
4    Relational     summaries Part II, first to     Ch. 6, 7, 8              Sec. 10.7, Ch. 11 (Finding 4)
     versus fuller materials    third
5    Consistency across the Part II, first to       Sec. 10.2, 10.3          Sec. 10.7, Ch. 11 (Findings 4
     six models                 third                                        and 6)
6    Preloading versus simul- Part II, first and    Ch. 7                    Sec. 10.7, Ch. 11 (Finding 3)
     taneous loading            second
7    Run-to-run stability       Part II, second     Sec. 7.6, 9.4            Sec. 10.8, Ch. 11 (Finding 6)
                                and fourth
8    Summary length             Part II, second     Ch. 7, Sec. 8.2          Sec. 10.7, Ch. 11 (Finding 4)
                                and third
9    Majority class, outcome Part II, first to      Sec. 10.4, 10.5          Sec. 10.7, Ch. 11 (Finding 6)
     class and cases            third
10   Model-specific      length Part II, second     Sec. 8.2                 Sec. 10.7, Ch. 11 (Finding 4)
     curves                     and third
11   Summarisation tool with Part II, fourth        Ch. 9                    Sec. 10.7, Ch. 11 (Finding 5)
     a fixed predictor

Thesis Table 10.14 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.