12 papers
PreScience: A Dataset and Benchmark for Scientific Forecasting
Anirudh Ajith, Amanpreet Singh, Jay DeYoung +7
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built a…
Artificial Intelligence Index Report 2026
Sha Sajadieh, Loredana Fattorini, Raymond Perrault +20
Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built around it can keep up. Governance frameworks…
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
Jonathan Bragg, Mike D'Arcy, Nishant Balepur +36
AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions o…
Omakase: proactive assistance with actionable suggestions for evolving scientific research projects
Pao Siangliulue, Jonathan Bragg, Doug Downey +2
As AI agents become increasingly capable of complex knowledge tasks, the lack of context limits their capability to proactively reason about a user's latent needs throughout a long…
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
Jena D. Hwang, Varsha Kishore, Amanpreet Singh +9
Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, alon…
Cocoa: Co-Planning and Co-Execution with AI Agents
K. J. Kevin Feng, Kevin Pu, Matt Latzke +6
As AI agents take on increasingly long-running tasks involving sophisticated planning and execution, there is a corresponding need for novel interaction designs that enable deeper…