4 papers
Language Models and Logic Programs for Trustworthy Tax Reasoning
William Jurayj, Nils Holzenberger, Benjamin Van Durme
According to the United States Internal Revenue Service, ``the average American spends and 13 hours filing their taxes''. Even beyond the U.S., tax filing requires complex…
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
Jiefu Ou, William Gantt Walden, Kate Sanders +13
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically genera…
Is That Your Final Answer? Test-Time Scaling Improves Selective Question Answering
William Jurayj, Jeffrey Cheng, Benjamin Van Durme
Scaling the test-time compute of large language models has demonstrated impressive performance on reasoning benchmarks. However, existing evaluations of test-time scaling make the…
Gaps or Hallucinations? Gazing into Machine-Generated Legal Analysis for Fine-grained Text Evaluations
Abe Bohan Hou, William Jurayj, Nils Holzenberger +2
Large Language Models (LLMs) show promise as a writing aid for professionals performing legal analyses. However, LLMs can often hallucinate in this setting, in ways difficult to re…