From the 1 of 8 linked papers with an AI index.
2 citations · 2 across the 6 of their papers we have counts for
6 papers · 1 filter
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21
The paper evaluates whether current AI agents can independently conduct open‑ended AI research by having them attempt to solve the central questions of two unpublished NeurIPS subm…
Life After Benchmark Saturation: A Case Study of CORE-Bench
Nitya Nadgir, Sayash Kapoor, Kangheng Liu +11
When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privileges accuracy and misses the opportunity t…
Towards a Science of AI Agent Reliability
Stephan Rabanser, Sayash Kapoor, Peter Kirgis +3
AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many agents still continue to fail in pr…
Open-World Evaluations for Measuring Frontier AI Capabilities
Sayash Kapoor, Peter Kirgis, Andrew Schwartz +15
Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be…
Log analysis is necessary for credible evaluation of AI agents
Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8
Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated or deflated by shortcuts and…
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…