From the 1 of 9 linked papers with an AI index.
4 citations · 4 across the 4 of their papers we have counts for
9 papers
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…
FLARE-AI: Flaw Reporting for AI
Shayne Longpre, Elaine Zhu, Carson Ezell +15
Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosystem is fragmented: researchers who identify…
Life After Benchmark Saturation: A Case Study of CORE-Bench
Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11
When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir +2
AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the development of useful agents, we need benchmark…
Towards a Science of AI Agent Reliability
Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in pr…
Open-World Evaluations for Measuring Frontier AI Capabilities
Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15
Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be…