5 papers
Can LLMs Identify Tax Abuse?
Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme
We investigate whether large language models can discover and analyze U.S. tax-minimization strategies. This real-world domain challenges even seasoned human experts, and progress…
LLMs Provide Unstable Answers to Legal Questions
Andrew Blair-Stanek, Benjamin Van Durme
An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when…
BLT: Can Large Language Models Handle Basic Legal Text?
Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme
We find that the best publicly available LLMs like GPT-4 and Claude currently perform poorly on basic legal text handling. This motivates the creation of a benchmark consisting of…
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
Abe Bohan Hou, William Jurayj, Nils Holzenberger +2
Large Language Models (LLMs) show promise as a writing aid for professionals performing legal analyses. However, LLMs can often hallucinate in this setting, in ways difficult to re…
CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation
Abe Bohan Hou, Orion Weller, Guanghui Qin +5
Legal professionals need to write analyses that rely on citations to relevant precedents, i.e., previous case decisions. Intelligent systems assisting legal professionals in writin…