Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models
arXiv:2401.01301 · doi:10.1093/jla/laae003
Abstract
Do large language models (LLMs) know the law? These models are increasingly being used to augment legal practice, education, and research, yet their revolutionary potential is threatened by the presence of hallucinations -- textual output that is not consistent with legal facts. We present the first systematic evidence of these hallucinations, documenting LLMs' varying performance across jurisdictions, courts, time periods, and cases. Our work makes four key contributions. First, we develop a typology of legal hallucinations, providing a conceptual framework for future research in this area. Second, we find that legal hallucinations are alarmingly prevalent, occurring between 58% of the time with ChatGPT 4 and 88% with Llama 2, when these models are asked specific, verifiable questions about random federal court cases. Third, we illustrate that LLMs often fail to correct a user's incorrect legal assumptions in a contra-factual question setup. Fourth, we provide evidence that LLMs cannot always predict, or do not always know, when they are producing legal hallucinations. Taken together, our findings caution against the rapid and unsupervised integration of popular LLMs into legal tasks. Even experienced lawyers must remain wary of legal hallucinations, and the risks are highest for those who stand to benefit from LLMs the most -- pro se litigants or those without access to traditional legal resources.
References in corpus (38)
- Survey of Hallucination in Natural Language Generation
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- On the Opportunities and Risks of Foundation Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
- Language Models (Mostly) Know What They Know
- Quantifying Memorization Across Neural Language Models
- Hallucination is Inevitable: An Innate Limitation of Large Language Models
- Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
- Algorithmic Monoculture and Social Welfare
- Towards Understanding Sycophancy in Language Models
- A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models
- Chatlaw: A Multi-Agent Legal Assistant based on a Role-Aligned Mixture-of-Experts Architecture
- Prompting GPT-3 To Be Reliable
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Factuality Enhanced Language Models for Open-Ended Text Generation
- Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs
- Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
- Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
- How well do LLMs cite relevant medical references? An evaluation framework and analyses
- Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
- Capturing Failures of Large Language Models via Human Cognitive Biases
- Simple synthetic data reduces sycophancy in large language models
- FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
- How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
- Lift Yourself Up: Retrieval-augmented Text Generation with Self Memory
- LawBench: Benchmarking Legal Knowledge of Large Language Models
- Explaining Legal Concepts with Augmented Large Language Models (GPT-4)
- Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
- Chain of Natural Language Inference for Reducing Large Language Model Ungrounded Hallucinations
- Fine-tuning Language Models for Factuality
- Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language Models
- Large Language Models in Cryptocurrency Securities Cases: Can a GPT Model Meaningfully Assist Lawyers?
- R-Tuning: Instructing Large Language Models to Say `I Don't Know'
- Comparing Hallucination Detection Metrics for Multilingual Generation
Cited by in corpus (11)
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
- Secret Use of Large Language Model (LLM)
- Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLM
- On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
- Natural Language Processing for the Legal Domain: A Survey of Tasks, Datasets, Models, and Challenges
- Explainable Rule Application via Structured Prompting: A Neural-Symbolic Approach
- Agentic AI and Hallucinations
- A Layered Multi-Expert Framework for Long-Context Mental Health Assessments
- Automated Theorem Provers Help Improve Large Language Model Reasoning
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?