9 papers
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
Delip Rao, Weiqiu You, Eric Wong +1
We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientif…
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
Muxin Liu, Delip Rao, Grace Kim +1
Scientific claim verification, the task of determining whether claims are entailed by scientific evidence, is fundamental to establishing discoveries in evidence while preventing m…
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
Delip Rao, Eric Wong, Chris Callison-Burch
Large language models and deep research agents supply citation URLs to support their claims, yet the reliability of these citations has not been systematically measured. We address…
BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation
Delip Rao, Chris Callison-Burch
Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive field-level errors stemming from omissio…
Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
Delip Rao, Chris Callison-Burch
Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduced to exact programmatic checks. Yet t…
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
Delip Rao, Chris Callison-Burch
Despite rapid progress in claim verification, we lack a systematic understanding of what reasoning these benchmarks actually exercise. We generate structured reasoning traces for 2…