CrossLaw/Part II · Nigeria/Stage 4 · Summarisation tool
14 / 26 · Part II · NigeriaThesis Chapter 9; Figure 9.1

With one predictor fixed, tools differed on five cases.

Stage 4 · Summarisation tool

Key takeaways

  • Tool accuracy ranged from 75.0% to 90.0% in run a.
  • Shorter summaries equaled or exceeded longer ones for each tool.
  • Gemini Notebook produced 10/11 usable summaries; BOFIA was missing.

What was tested & why

Stage 4 · Summarisation tool

ChatGPT predicted from stored summaries prepared by five different tools. Claude’s 100–150-word summaries produced the highest two-run mean (92.5%). Tool differences arose from five cases and are not statistically separable on this sample.

Evidence boundary

The comparison with the stage-2 full-material result crosses stages and is descriptive only.

Source: Thesis Chapter 9; Figure 9.1 · source-reported unless otherwise noted.

Results / visual evidence

Data visualization · source-reported

Fixed ChatGPT predictor, summaries by tool

100–150 words150–300 words— 70% always-Dismissed reference
ChatGPT
90.0%
85.0%
Claude
90.0%
85.0%
Microsoft Copilot
75.0%
75.0%
Gemini
75.0%
75.0%
Gemini Notebook
80.0%
80.0%

Thesis Tables 9.2–9.3 and Figure 9.1 · 20 Nigerian cases per tool and length. Gemini Notebook produced 10/11 source summaries.

Source tables / thesis transcription

Inspect the evidence

Scroll wide tables horizontally. Source notes and qualifications remain with their tables.

Table 9.1 · source-reported

Stored summaries by tool: summaries produced and word-range compliance (summary text only, excluding the title)

                                           100–150 words              150–300 words

                               Summaries   In     Mean words          In    Mean words
           Tool
                                produced range     (range)          range    (range)

           ChatGPT               11/11   11/11    138.5 (127–146)   11/11   223.7 (207–246)
           Claude                11/11   11/11    148.4 (145–150)   11/11   295.0 (287–299)
           Microsoft Copilot     11/11   11/11    133.9 (117–145)   11/11   189.4 (164–209)
           Gemini                11/11   11/11    141.4 (136–146)   11/11   274.4 (253–289)
           Gemini Notebook       10/11   10/10    140.0 (127–148)   10/10   182.7 (155–231)

 All five tools were supplied with the same 11 sources. Gemini Notebook did not produce a
 usable summary of the Banks and Other Financial Institutions Act 2020 (BOFIA) at either
  length, so its in-range counts and word statistics are over its 10 usable summaries. Word
                         counts by source are given in Appendix 12.4.

Thesis Table 9.1 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 9.2 · source-reported

ChatGPT prediction results by summarisation tool at 100–150 words per source (run a)

                                      Dism. Allow.     Bal.   Mean
Summariser          Correct    Acc.                                           Brier       EV      SF
                                      (/14)   (/6)     acc.   conf. Conf.−
                                                                      Acc.

ChatGPT                 18    90.0%      12       6   92.9%   60.0%   −30.1   0.1861   +0.4855     2
Claude                  18    90.0%      12       6   92.9%   60.5%   −29.6   0.1849   +0.4875     2
Microsoft Copilot       15    75.0%      10       5   77.4%   63.9%   −11.2   0.2024   +0.3325     4
Gemini                  15    75.0%      11       4   72.6%   61.1%   −14.0   0.2371   +0.2875     5
Gemini Notebook         16    80.0%      11       5   81.0%   59.7%   −20.4   0.1981   +0.3695     4

Pooled              82/100    82.0%   56/70   26/30   83.3%   61.0%   −21.0 0.2017 +0.3925        17

                        Pooled Wilson 95% interval for 82/100: 73.3–88.3%.

Thesis Table 9.2 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 9.3 · source-reported

ChatGPT prediction results by summarisation tool at 150–300 words per source (run a)

                                      Dism. Allow.     Bal.   Mean
Summariser          Correct    Acc.                                           Brier       EV      SF
                                      (/14)   (/6)     acc.   conf. Conf.−
                                                                      Acc.

ChatGPT                 17    85.0%      12       5   84.5%   63.2%   −21.8   0.2163   +0.4190     3
Claude                  17    85.0%      12       5   84.5%   64.2%   −20.8   0.2071   +0.4270     3
Microsoft Copilot       15    75.0%      10       5   77.4%   66.1%    −8.9   0.2044   +0.3400     4
Gemini                  15    75.0%      11       4   72.6%   63.9%   −11.1   0.2358   +0.2990     5
Gemini Notebook         16    80.0%      11       5   81.0%   61.6%   −18.5   0.1943   +0.3785     4

Pooled              80/100    80.0%   56/70   24/30   80.0%   63.8%   −16.2 0.2116 +0.3727        19

Pooled Wilson 95% interval for 80/100: 71.1–86.7%. The pooled mean confidence, Brier score
                and EV are unweighted means of the five tool conditions.

Thesis Table 9.3 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 9.4 · source-reported

Case-level results by summarisation tool (run a): number of tools (out of 5) whose summaries led to the correct verdict, and the tools whose summaries led to an error

Case     GT       100–150   Wrong at 100–150             150–300   Wrong at 150–300

NG 001   Dism.      5/5     –                              5/5     –
NG 002   Dism.      2/5     Copilot, Gemini, Gemini        2/5     Copilot, Gemini, Gemini
                            Notebook                               Notebook
NG 003   Dism.      5/5     –                              5/5     –
NG 004   Dism.      5/5     –                              5/5     –
NG 005   Dism.      5/5     –                              5/5     –
NG 006   Allow.     5/5     –                              5/5     –
NG 007   Dism.      5/5     –                              5/5     –
NG 008   Dism.      5/5     –                              5/5     –
NG 009   Dism.      0/5     ChatGPT, Claude, Copilot,      0/5     ChatGPT, Claude, Copilot,
                            Gemini, Gemini Notebook                Gemini, Gemini Notebook
NG 010   Dism.      5/5     –                              5/5     –
NG 011   Allow.     4/5     Gemini                         2/5     ChatGPT, Claude, Gemini
NG 012   Dism.      5/5     –                              5/5     –
NG 013   Allow.     5/5     –                              5/5     –
NG 014   Allow.     5/5     –                              5/5     –
NG 015   Dism.      5/5     –                              5/5     –
NG 016   Dism.      5/5     –                              5/5     –
NG 017   Allow.     4/5     Copilot                        4/5     Copilot
NG 018   Dism.      4/5     Copilot                        4/5     Copilot
NG 019   Allow.     3/5     Gemini, Gemini Notebook        3/5     Gemini, Gemini Notebook
NG 020   Dism.      0/5     ChatGPT, Claude, Copilot,      0/5     ChatGPT, Claude, Copilot,
                            Gemini, Gemini Notebook                Gemini, Gemini Notebook

 Copilot = Microsoft Copilot. Full verdicts and confidence scores are given in Appendix 12.3.

Thesis Table 9.4 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 9.5 · source-reported

100–150 against 150–300 words per source, by summarisation tool (run a)

Summariser                  100–150             150–300        Change (pp)       Same verdict at
                                                                                   both lengths

ChatGPT                       90.0%                85.0%               −5.0                 19/20
Claude                        90.0%                85.0%               −5.0                 19/20
Microsoft Copilot             75.0%                75.0%                0.0                 20/20
Gemini                        75.0%                75.0%                0.0                 20/20
Gemini Notebook               80.0%                80.0%                0.0                 20/20

Pooled                       82.0%                80.0%                −2.0               98/100

Thesis Table 9.5 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.

Table 9.6 · source-reported

Run a against run b for the ChatGPT- and Claude-summary con- ditions

                                       Mean   Same     Run b       Run b
Summariser,         Run a Run b                                          Verdicts that changed
                                    (2 runs) verdict bal. acc. conf.<50%
length                                                                   (a→b, run-b result)

ChatGPT, 100–150    90.0%   85.0%     87.5%   19/20    89.3%        0/20   NG 002: D→A (wrong)
Claude, 100–150     90.0%   95.0%     92.5%   19/20    96.4%       14/20   NG 009: A→D (right)
ChatGPT, 150–300    85.0%   80.0%     82.5%   19/20    81.0%        0/20   NG 002: D→A (wrong)
Claude, 150–300     85.0%   85.0%     85.0%   18/20    84.5%       10/20   NG 005: D→A (wrong),
                                                                           NG 009: A→D (right)

 Run-b wrong cases: ChatGPT summaries, NG 002, 009 and 020 (100–150) and NG 002, 009,
   011 and 020 (150–300). Claude summaries, NG 020 (100–150) and NG 005, 011 and 020
                   (150–300). Run-b verdicts are given in Appendix 12.3.

Thesis Table 9.6 · source transcription; layout preserved for multi-line headings and notes. Verify against the source PDF for formal citation.