15 citations · 15 across the 2 of their papers we have counts for
4 papers
HCAST: Human-Calibrated Autonomy Software Tasks
David Rein, Joel Becker, Amy Deng +19
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker +23
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Hjalmar Wijk, Tao Lin, Joel Becker +20
Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du +10
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refe…