Key takeaways
- Tool accuracy ranged from 75.0% to 90.0% in run a.
- Shorter summaries equaled or exceeded longer ones for each tool.
- Gemini Notebook produced 10/11 usable summaries; BOFIA was missing.
What was tested & why
Stage 4 · Summarisation tool
ChatGPT predicted from stored summaries prepared by five different tools. Claude’s 100–150-word summaries produced the highest two-run mean (92.5%). Tool differences arose from five cases and are not statistically separable on this sample.
The comparison with the stage-2 full-material result crosses stages and is descriptive only.
Source: Thesis Chapter 9; Figure 9.1 · source-reported unless otherwise noted.
Results / visual evidence
Fixed ChatGPT predictor, summaries by tool
Thesis Tables 9.2–9.3 and Figure 9.1 · 20 Nigerian cases per tool and length. Gemini Notebook produced 10/11 source summaries.
Source tables / thesis transcription
Inspect the evidence
Scroll wide tables horizontally. Source notes and qualifications remain with their tables.
Stored summaries by tool: summaries produced and word-range compliance (summary text only, excluding the title)
100–150 words 150–300 words
Summaries In Mean words In Mean words
Tool
produced range (range) range (range)
ChatGPT 11/11 11/11 138.5 (127–146) 11/11 223.7 (207–246)
Claude 11/11 11/11 148.4 (145–150) 11/11 295.0 (287–299)
Microsoft Copilot 11/11 11/11 133.9 (117–145) 11/11 189.4 (164–209)
Gemini 11/11 11/11 141.4 (136–146) 11/11 274.4 (253–289)
Gemini Notebook 10/11 10/10 140.0 (127–148) 10/10 182.7 (155–231)
All five tools were supplied with the same 11 sources. Gemini Notebook did not produce a
usable summary of the Banks and Other Financial Institutions Act 2020 (BOFIA) at either
length, so its in-range counts and word statistics are over its 10 usable summaries. Word
counts by source are given in Appendix 12.4.Thesis Table 9.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
ChatGPT prediction results by summarisation tool at 100–150 words per source (run a)
Dism. Allow. Bal. Mean
Summariser Correct Acc. Brier EV SF
(/14) (/6) acc. conf. Conf.−
Acc.
ChatGPT 18 90.0% 12 6 92.9% 60.0% −30.1 0.1861 +0.4855 2
Claude 18 90.0% 12 6 92.9% 60.5% −29.6 0.1849 +0.4875 2
Microsoft Copilot 15 75.0% 10 5 77.4% 63.9% −11.2 0.2024 +0.3325 4
Gemini 15 75.0% 11 4 72.6% 61.1% −14.0 0.2371 +0.2875 5
Gemini Notebook 16 80.0% 11 5 81.0% 59.7% −20.4 0.1981 +0.3695 4
Pooled 82/100 82.0% 56/70 26/30 83.3% 61.0% −21.0 0.2017 +0.3925 17
Pooled Wilson 95% interval for 82/100: 73.3–88.3%.Thesis Table 9.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
ChatGPT prediction results by summarisation tool at 150–300 words per source (run a)
Dism. Allow. Bal. Mean
Summariser Correct Acc. Brier EV SF
(/14) (/6) acc. conf. Conf.−
Acc.
ChatGPT 17 85.0% 12 5 84.5% 63.2% −21.8 0.2163 +0.4190 3
Claude 17 85.0% 12 5 84.5% 64.2% −20.8 0.2071 +0.4270 3
Microsoft Copilot 15 75.0% 10 5 77.4% 66.1% −8.9 0.2044 +0.3400 4
Gemini 15 75.0% 11 4 72.6% 63.9% −11.1 0.2358 +0.2990 5
Gemini Notebook 16 80.0% 11 5 81.0% 61.6% −18.5 0.1943 +0.3785 4
Pooled 80/100 80.0% 56/70 24/30 80.0% 63.8% −16.2 0.2116 +0.3727 19
Pooled Wilson 95% interval for 80/100: 71.1–86.7%. The pooled mean confidence, Brier score
and EV are unweighted means of the five tool conditions.Thesis Table 9.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Case-level results by summarisation tool (run a): number of tools (out of 5) whose summaries led to the correct verdict, and the tools whose summaries led to an error
Case GT 100–150 Wrong at 100–150 150–300 Wrong at 150–300
NG 001 Dism. 5/5 – 5/5 –
NG 002 Dism. 2/5 Copilot, Gemini, Gemini 2/5 Copilot, Gemini, Gemini
Notebook Notebook
NG 003 Dism. 5/5 – 5/5 –
NG 004 Dism. 5/5 – 5/5 –
NG 005 Dism. 5/5 – 5/5 –
NG 006 Allow. 5/5 – 5/5 –
NG 007 Dism. 5/5 – 5/5 –
NG 008 Dism. 5/5 – 5/5 –
NG 009 Dism. 0/5 ChatGPT, Claude, Copilot, 0/5 ChatGPT, Claude, Copilot,
Gemini, Gemini Notebook Gemini, Gemini Notebook
NG 010 Dism. 5/5 – 5/5 –
NG 011 Allow. 4/5 Gemini 2/5 ChatGPT, Claude, Gemini
NG 012 Dism. 5/5 – 5/5 –
NG 013 Allow. 5/5 – 5/5 –
NG 014 Allow. 5/5 – 5/5 –
NG 015 Dism. 5/5 – 5/5 –
NG 016 Dism. 5/5 – 5/5 –
NG 017 Allow. 4/5 Copilot 4/5 Copilot
NG 018 Dism. 4/5 Copilot 4/5 Copilot
NG 019 Allow. 3/5 Gemini, Gemini Notebook 3/5 Gemini, Gemini Notebook
NG 020 Dism. 0/5 ChatGPT, Claude, Copilot, 0/5 ChatGPT, Claude, Copilot,
Gemini, Gemini Notebook Gemini, Gemini Notebook
Copilot = Microsoft Copilot. Full verdicts and confidence scores are given in Appendix 12.3.Thesis Table 9.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
100–150 against 150–300 words per source, by summarisation tool (run a)
Summariser 100–150 150–300 Change (pp) Same verdict at
both lengths
ChatGPT 90.0% 85.0% −5.0 19/20
Claude 90.0% 85.0% −5.0 19/20
Microsoft Copilot 75.0% 75.0% 0.0 20/20
Gemini 75.0% 75.0% 0.0 20/20
Gemini Notebook 80.0% 80.0% 0.0 20/20
Pooled 82.0% 80.0% −2.0 98/100Thesis Table 9.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.
Run a against run b for the ChatGPT- and Claude-summary con- ditions
Mean Same Run b Run b
Summariser, Run a Run b Verdicts that changed
(2 runs) verdict bal. acc. conf.<50%
length (a→b, run-b result)
ChatGPT, 100–150 90.0% 85.0% 87.5% 19/20 89.3% 0/20 NG 002: D→A (wrong)
Claude, 100–150 90.0% 95.0% 92.5% 19/20 96.4% 14/20 NG 009: A→D (right)
ChatGPT, 150–300 85.0% 80.0% 82.5% 19/20 81.0% 0/20 NG 002: D→A (wrong)
Claude, 150–300 85.0% 85.0% 85.0% 18/20 84.5% 10/20 NG 005: D→A (wrong),
NG 009: A→D (right)
Run-b wrong cases: ChatGPT summaries, NG 002, 009 and 020 (100–150) and NG 002, 009,
011 and 020 (150–300). Claude summaries, NG 020 (100–150) and NG 005, 011 and 020
(150–300). Run-b verdicts are given in Appendix 12.3.Thesis Table 9.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.