activity
20242026
collaborators

5 papers

q-fin.CP2026

Can LLMs Identify Tax Abuse?

Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme

We investigate whether large language models can discover and analyze U.S. tax-minimization strategies. This real-world domain challenges even seasoned human experts, and progress…

cs.CL2025

LLMs Provide Unstable Answers to Legal Questions

Andrew Blair-Stanek, Benjamin Van Durme

An LLM is stable if it reaches the same conclusion when asked the identical question multiple times. We find leading LLMs like gpt-4o, claude-3.5, and gemini-1.5 are unstable when…

cs.CL2024

BLT: Can Large Language Models Handle Basic Legal Text?

Andrew Blair-Stanek, Nils Holzenberger, Benjamin Van Durme

We find that the best publicly available LLMs like GPT-4 and Claude currently perform poorly on basic legal text handling. This motivates the creation of a benchmark consisting of…

cs.CL2024

Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations

Abe Bohan Hou, William Jurayj, Nils Holzenberger +2

Large Language Models (LLMs) show promise as a writing aid for professionals performing legal analyses. However, LLMs can often hallucinate in this setting, in ways difficult to re…

cs.CL2024

CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation

Abe Bohan Hou, Orion Weller, Guanghui Qin +5

Legal professionals need to write analyses that rely on citations to relevant precedents, i.e., previous case decisions. Intelligent systems assisting legal professionals in writin…