7 papers
PreScience: A Dataset and Benchmark for Scientific Forecasting
Anirudh Ajith, Amanpreet Singh, Jay DeYoung +7
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built a…
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
Nishant Balepur, Malachi Hamada, Varsha Kishore +9
Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols…
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Jonathan Bragg, Mike D'Arcy, Nishant Balepur +36
AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions o…
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
Pao Siangliulue, Jonathan Bragg, Doug Downey +2
As AI agents become increasingly capable of complex knowledge tasks, the lack of context limits their capability to proactively reason about a user's latent needs throughout a long…
Generating Literature-Driven Scientific Theories at Scale
Peter Jansen, Peter Clark, Doug Downey +1
Contemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theor…
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
Yilun Zhao, Kaiyan Zhang, Tiansheng Hu +15
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific liter…