CrossLaw/Evidence library/References
25 / 26 · Evidence libraryThesis bibliography

The submitted thesis provides the bibliography.

References

Key takeaways

  • Bibliographic entries are transcribed from the thesis.
  • References support context; results derive from the CrossLaw experiments.
  • Some PDF-extracted URLs contain spaces and should be checked against the original PDF.

What was tested & why

References

The reference list below is source-reported from the submitted thesis. Extraction preserves citation text while line breaks may differ from the printed PDF.

Source: Thesis bibliography · source-reported unless otherwise noted.

Thesis bibliography

Bibliography

[1] Akhihiero, P. A. (2025). The new frontiers in artificial intelligence in the
    legal profession in Nigeria. High Court of Edo State, Nigeria. https :
    / / edojudiciary . gov . ng / wp - content / uploads / 2025 / 10 / THE - NEW -
    FRONTIERS- IN- ARTIFICIAL- INTELLIGENCE- IN- THE- LEGAL- PROFESSION-
    IN-NIGERIA.pdf

[2] Ariai, F., Mackenzie, J., & Demartini, G. (2025). Natural language processing
    for the legal domain: A survey of tasks, datasets, models, and challenges.
    ACM Computing Surveys, 58 (6), Article 163, 1–37. https://doi.org/10.
    1145/3777009

[3] Australian Government, Department of Industry, Science and Resources.
    (2025). National AI Plan. https://www.industry.gov.au/publications/
    national-ai-plan

[4] Australian Government, Digital Transformation Agency. (2025). AI Plan for
    the Australian Public Service 2025. https://www.digital.gov.au/ai-
    plan-australian-public-service-2025-appendix-plan-deliverables

[5] Ayana, G., Dese, K., Daba, H., Mellado, B., Badu, K., Yamba, E., Faye,
    S., Ondua, M., Nsagha, D., Nkweteyim, D., & Kong, J. (2024). Decolonizing
    global AI governance: Assessment of the state of decolonized AI governance
    in Sub-Saharan Africa. Royal Society Open Science, 11 (6), 231994. https:
    //doi.org/10.1098/rsos.231994

[6] BAM & GAD Solicitors. (2024, April 23). Artificial intelligence and lawyers
    in Nigeria. https : / / bamandgadsolicitors . com . ng / 2024 / 04 / 23 /
    artificial-intelligence-and-lawyers-in-nigeria/

[7] Bednar, N., Cleveland, D. R., Erbsen, A., & Schwarcz, D. (2026). Artificial
    intelligence and human legal reasoning (Minnesota Legal Studies Research
    Paper No. 2026-21). SSRN. https://doi.org/10.2139/ssrn.6525800
                                      148
BIBLIOGRAPHY                                                                     149

 [8] Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021).
     On the dangers of stochastic parrots: Can language models be too big? In
     Proceedings of the 2021 ACM Conference on Fairness, Accountability, and
     Transparency, 610–623. https://doi.org/10.1145/3442188.3445922

 [9] Braun, C., Lilienbeck, A., & Mentjukov, D. (2025). The hidden structure: Im-
     proving legal document understanding through explicit text formatting. arXiv
     preprint arXiv:2505.12837. https://doi.org/10.48550/arXiv.2505.12837

[10] Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I.,
     Katz, D. M., & Aletras, N. (2022). LexGLUE: A benchmark dataset for legal
     language understanding in English. In Proceedings of the 60th Annual Meeting
     of the Association for Computational Linguistics, 4310–4330. https://doi.
     org/10.18653/v1/2022.acl-long.297

[11] Cochran, W. G. (1950). The comparison of percentages in matched samples.
     Biometrika, 37 (3/4), 256–266.

[12] Cohen, I. G., Babic, B., Gerke, S., Xia, Q., Evgeniou, T., & Wertenbroch,
     K. (2023). How AI can learn from the law: Putting humans in the loop only
     on appeal. npj Digital Medicine, 6, Article 160. https://doi.org/10.1038/
     s41746-023-00906-8

[13] Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd
     ed.). Lawrence Erlbaum Associates.

[14] Council of Europe. (2024). Council of Europe Framework Convention on Arti-
     ficial Intelligence and Human Rights, Democracy and the Rule of Law. Council
     of Europe Treaty Series, No. 225. https://rm.coe.int/1680afae3c

[15] Curran, D., Sporne, V., Frermann, L., & Paterson, J. (2025). Place mat-
     ters: Comparing LLM hallucination rates for place-based legal queries. arXiv
     preprint arXiv:2511.06700. https://doi.org/10.48550/arXiv.2511.06700

[16] Dahl, M., Magesh, V., Suzgun, M., & Ho, D. E. (2024). Large legal fictions:
     Profiling legal hallucinations in large language models. Journal of Legal Anal-
     ysis, 16 (1), 64–93. https://doi.org/10.1093/jla/laae003

[17] Dalzell, T., Smaill, A., & Koay, J. (2025). Analysis of large language model
     prompting using Toulmin’s model of argumentation. In AI 2025: Advances
     in Artificial Intelligence, 38th Australasian Joint Conference on Artificial In-
     telligence. Springer. (In press)
BIBLIOGRAPHY                                                                    150

[18] Deroy, A., et al. (2024). Applicability of large language models and generative
     models for legal case judgement summarization. Artificial Intelligence and
     Law, 33, 1007–1050. https://doi.org/10.1007/s10506-024-09411-z

[19] Dor, L., & Coglianese, C. (2021). Procurement as AI governance. IEEE Trans-
     actions on Technology and Society, 2 (4), 192–199. https://doi.org/10.
     1109/TTS.2021.3111764

[20] El Hamdani, R., Bonald, T., Malliaros, F. D., Holzenberger, N., & Suchanek,
     F. (2024). The Factuality of Large Language Models in the Legal Domain.
     Proceedings of the 33rd ACM International Conference on Information and
     Knowledge Management. https://doi.org/10.1145/3627673.3679961

[21] European Parliament and Council of the European Union. (2024). Regulation
     (EU) 2024/1689 (Artificial Intelligence Act). Official Journal of the European
     Union, L 2024/1689. http://data.europa.eu/eli/reg/2024/1689/oj

[22] Everitt, B. S. (1998). The Cambridge dictionary of statistics. Cambridge Uni-
     versity Press.

[23] Farzi, N., et al. (2026). Supporting Humans in Evaluating AI Summaries of
     Legal Depositions. Proceedings of the 2026 Conference on Human Information
     Interaction and Retrieval. https://doi.org/10.1145/3786304.3787923

[24] Federal Ministry of Communications, Innovation and Digital Economy.
     (2025). National Artificial Intelligence Strategy (NAIS) 2025–2029. https:
     / / ncair . nitda . gov . ng / wp - content / uploads / 2025 / 09 / National -
     Artificial-Intelligence-Strategy-19092025.pdf

[25] Fei, Z., Shen, X., Zhu, D., Zhou, F., Han, Z., Huang, A., Zhang, S., Chen,
     K., Yin, Z., Shen, Z., Ge, J., & Ng, V. (2024). LawBench: Benchmarking
     legal knowledge of large language models. In Proceedings of EMNLP 2024,
     7933–7962. https://doi.org/10.18653/v1/2024.emnlp-main.452

[26] Finextra Research. (2024, August 6). How much does it cost to build an
     AI system? Finextra Community Blogs. https : / / www . finextra . com /
     blogposting/26576/how-much-does-it-cost-to-build-an-ai-system

[27] Fricker, M. (2007). Epistemic injustice: Power and the ethics of know-
     ing. Oxford University Press. https://doi.org/10.1093/acprof:oso/
     9780198237907.001.0001
BIBLIOGRAPHY                                                                      151

[28] Gneiting, T., & Raftery, A. E. (2007). Strictly proper scoring rules, prediction,
     and estimation. Journal of the American Statistical Association, 102 (477),
     359–378. https://doi.org/10.1198/016214506000001437

[29] Guha, N., Nyarko, J., Ho, D. E., Ré, C., Chilton, A., Narayana, A., et al.
     (2023). LegalBench: A collaboratively built benchmark for measuring legal
     reasoning in large language models (Osgoode Legal Studies Research Paper
     No. 4583531). SSRN. https://doi.org/10.2139/ssrn.4583531

[30] Han, J., Burgess, P., & Shareghi, E. (2026). Legal citation prediction
     with LLMs: A comparative evaluation of instruction tuning, retrieval, and
     jurisdiction-specific pre-training on the AusLaw citation benchmark. Artifi-
     cial Intelligence and Law. https://doi.org/10.1007/s10506-026-09506-9

[31] Handa & Mallick [2024] FedCFamC2F 957 (Federal Circuit and Family Court
     of Australia, Division 2).

[32] Hemrajani, R. (2025). Evaluating the role of large language models in legal
     practice in India. arXiv preprint arXiv:2508.09713. https://doi.org/10.
     48550/arXiv.2508.09713

[33] Hu, Y., Liu, H., Wang, C., Li, K., Wu, T., Li, H., et al. (2026). Evaluation of
     large language models in legal applications: Challenges, methods, and future
     directions. arXiv preprint arXiv:2601.15267. https://doi.org/10.48550/
     arXiv.2601.15267

[34] International Organization for Standardization. (2023). ISO/IEC 42001:2023
     Information technology, Artificial intelligence, Management system. ISO.
     https://www.iso.org/standard/81230.html

[35] Ioannou, A., Shiamishis, A., Hollenstein, N., & Gurel, N. (2025). Evaluating
     the limits of large language models in multilingual legal reasoning. arXiv
     preprint arXiv:2509.22472. https://doi.org/10.48550/arXiv.2509.22472

[36] Jonckheere, A. R. (1954). A distribution-free k-sample test against ordered
     alternatives. Biometrika, 41 (1/2), 133–145.

[37] Klem, S., & Moubayed, N. (2025). LLMs for LLMs: A structured prompt-
     ing methodology for long legal documents. arXiv preprint arXiv:2509.02241.
     https://doi.org/10.48550/arXiv.2509.02241
BIBLIOGRAPHY                                                                    152

[38] Li, H., Chen, Y., Ai, Q., Wu, Y., Zhang, R., & Liu, Y. (2024). LexEval: A
     comprehensive Chinese legal benchmark for evaluating large language mod-
     els. arXiv preprint arXiv:2409.20288. https://doi.org/10.48550/arXiv.
     2409.20288

[39] Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F.,
     & Liang, P. (2024). Lost in the Middle: How Language Models Use Long
     Contexts. Transactions of the Association for Computational Linguistics, 12,
     157–173. https://doi.org/10.1162/tacl_a_00638

[40] Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D.
     E. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal
     Research Tools. arXiv preprint arXiv:2405.20362. https://doi.org/10.
     48550/arxiv.2405.20362

[41] Martin, L., Whitehouse, N., Yiu, S., Catterson, L., & Perera, R. (2024).
     Better call GPT: Comparing large language models against lawyers. arXiv
     preprint arXiv:2401.16212. https://doi.org/10.48550/arXiv.2401.16212

[42] Mata v. Avianca, Inc., No. 1:22-cv-01461 (S.D.N.Y. June 22, 2023).

[43] Mavundla v MEC: Department of Co-Operative Government and Traditional
     Affairs KwaZulu-Natal and Others (7940/2024P) [2025] ZAKZPHC 2 (High
     Court of South Africa, KwaZulu-Natal Local Division, Pietermaritzburg).

[44] McNemar, Q. (1947). Note on the sampling error of the difference between
     correlated proportions or percentages. Psychometrika, 12 (2), 153–157.

[45] Mentzingen, H., António, N., & Bação, F. (2025). Effectiveness in retriev-
     ing legal precedents: Exploring text summarization and cutting-edge lan-
     guage models toward a cost-efficient approach. Artificial Intelligence and Law.
     https://doi.org/10.1007/s10506-025-09440-2

[46] Moskvichev, A., & Sejdinovic, D. (2025). All models are miscalibrated, but
     some less so: A theoretical and empirical comparison of calibration metrics.
     In AI 2025: Advances in Artificial Intelligence, 38th Australasian Joint Con-
     ference on Artificial Intelligence. Springer. (In press)

[47] Mutswiri, P., Hapanyengwi, G., & Chipfumbu, C. (2025). Artificial intelli-
     gence (AI) ethics meets Ubuntu: Towards a context-aware governance model
     for sustainable innovation in Africa. International Journal of Scientific Re-
     search and Management, 13 (8), EC02. https://doi.org/10.18535/ijsrm/
     v13i08.ec02
BIBLIOGRAPHY                                                                    153

[48] National Institute of Standards and Technology. (2023). Artificial intelligence
     risk management framework (AI RMF 1.0) (NIST AI 100-1). U.S. Depart-
     ment of Commerce. https://doi.org/10.6028/NIST.AI.100-1

[49] Nazly, M., & Yapa, P. (2026). Legal document summarization: a short review.
     Frontiers in Artificial Intelligence, 9. https://doi.org/10.3389/frai.
     2026.1787315

[50] Nigam, S., et al. (2025). NyayaRAG: Realistic Legal Judgment Predic-
     tion with RAG under the Indian Common Law System. arXiv preprint
     arXiv:2508.00709. https://doi.org/10.48550/arxiv.2508.00709

[51] Obinna, A., & Kess-Momoh, J. (2024). Developing a conceptual technical
     framework for ethical AI in procurement with emphasis on legal oversight.
     GSC Advanced Research and Reviews, 19 (1), 150–162. https://doi.org/
     10.30574/gscarr.2024.19.1.0149

[52] Oluka, A. (2024). Mitigating biases in training data: Technical and legal
     challenges for Sub-Saharan Africa. International Journal of Applied Research
     in Business and Management, 5 (1), 1–15. https://doi.org/10.51137/
     ijarbm.2024.5.1.10

[53] Organisation for Economic Co-operation and Development. (2024). OECD
     AI principles. OECD AI Policy Observatory. https://oecd.ai/en/ai-
     principles

[54] Parizi, A., Liu, Y., Nokku, P., Gholamian, S., & Emerson, D. (2023). A com-
     parative study of prompting strategies for legal text classification. In Pro-
     ceedings of the Natural Legal Language Processing Workshop 2023, 245–256.
     https://doi.org/10.18653/v1/2023.nllp-1.25

[55] Pathak, D., Kumar, H., Roy, A., George, F., Verma, M., & Moogi, P. (2025).
     Detecting silent failures in multi-agentic AI trajectories. arXiv preprint
     arXiv:2511.04032. https://arxiv.org/abs/2511.04032

[56] Potts, C., & Sudhof, M. (2026). Invisible failures in human-AI interactions.
     arXiv preprint arXiv:2603.15423. https : / / doi . org / 10 . 48550 / arXiv .
     2603.15423

[57] Roohi, A., van der Weel, M., & van Deemter, K. (2024). Beyond factualism:
     Evaluating LLM calibration on non-factual tasks. In AI 2024: Advances in
BIBLIOGRAPHY                                                                     154

    Artificial Intelligence, 37th Australasian Joint Conference on Artificial In-
    telligence (LNAI Vol. 15441, pp. 89–104). Springer. https://doi.org/10.
    1007/978-981-96-0345-5_7

[58] Santosh, T.Y.S.S., Weiss, C., & Grabmair, M. (2024). LexSumm and LexT5:
     Benchmarking and Modeling Legal Summarization Tasks in English. In Pro-
     ceedings of the Natural Legal Language Processing Workshop 2024, 381–403.
     https://doi.org/10.18653/v1/2024.nllp-1.35

[59] Sheskin, D. J. (2011). Handbook of parametric and nonparametric statistical
     procedures (5th ed.). Chapman and Hall/CRC. https://doi.org/10.1201/
     9780429186196

[60] Shui, R., et al. (2023). A Comprehensive Evaluation of Large Language
     Models on Legal Judgment Prediction. arXiv preprint arXiv:2310.11761.
     https://doi.org/10.48550/arxiv.2310.11761

[61] Song, Y., Qin, Y., Huang, R., Chen, Y., & Lin, C. (2025). Legal text
     summarization via judicial syllogism with large language models. Journal of
     King Saud University Computer and Information Sciences, 37, 111. https:
     //doi.org/10.1007/s44443-025-00113-3

[62] Stiel, M., Chen, T., & Mersiades, P. (2024, May 30). The Allens AI Australian
     law benchmark 2024: How proficient are generative AI tools in emulating the
     role of a human lawyer providing Australian legal advice? Allens. https:
     //www.allens.com.au/insights-news/explore/2024/the-allens-ai-
     australian-law-benchmark/

[63] Stiel, M., Chen, T., & Tridgell, A. (2025, June 26). Breakthroughs and pitfalls:
     The 2025 Allens AI Australian law benchmark. Allens. https://www.allens.
     com.au/insights-news/explore/2025/the-allens-ai-australian-law-
     benchmark-2025/

[64] Su, W., et al. (2025). JuDGE: Benchmarking Judgment Document Generation
     for Chinese Legal System. Proceedings of the 48th International ACM SIGIR
     Conference on Research and Development in Information Retrieval. https:
     //doi.org/10.1145/3726302.3730295

[65] UNESCO. (2021). Recommendation on the ethics of artificial intelligence.
     United Nations Educational, Scientific and Cultural Organization. https:
     //unesdoc.unesco.org/ark:/48223/pf0000381137
BIBLIOGRAPHY                                                                 155

[66] Wahidur, R. S. M., Kim, S., Choi, H., Bhatti, D. S., & Lee, H. (2025). Legal
     query RAG. IEEE Access, 13, 36978–36994. https://doi.org/10.1109/
     ACCESS.2025.3542125

[67] Wang, X., Kim, H., Rahman, S., Mitra, K., & Miao, Z. (2024). Human-
     LLM collaborative annotation through effective verification of LLM labels.
     In Proceedings of the 2024 CHI Conference on Human Factors in Computing
     Systems, Article 303, 1–21. https://doi.org/10.1145/3613904.3641960

[68] Wu, Y., et al. (2023). Precedent-Enhanced Legal Judgment Prediction with
     LLM and Domain-Model Collaboration. arXiv preprint arXiv:2310.09241.
     https://doi.org/10.48550/arxiv.2310.09241

[69] Zambrano, G. (2024). Case law as data: Prompt engineering strategies for
     case outcome extraction with large language models in a zero-shot setting.
     Law, Technology and Humans, 6 (2), 45–62. https://doi.org/10.5204/
     lthj.3623

[70] Zhang, L., Savelka, J., & Ashley, K. (2025). Do LLMs Truly “Understand”
     When a Precedent Is Overruled? Legal Knowledge and Information Systems.
     https://doi.org/10.3233/FAIA251592