most citedMeasuring AI Ability to Complete Long Software Tasks

14 citations · 14 across the 1 of their papers we have counts for

collaborators

5 papers

cs.AI202614 cited

Measuring AI Ability to Complete Long Software Tasks

Thomas Kwa, Ben West, Joel Becker +23

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…

cs.CY2025

Forecasting AI Time Horizon Under Compute Slowdowns

Parker Whitfill, Ben Snodin, Joel Becker

METR's time horizon metric has grown exponentially since 2019, along with compute. However, it is unclear whether compute scaling will persist at current rates through 2030, raisin…

cs.AI2025

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Joel Becker, Nate Rush, Elizabeth Barnes +1

Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI to…

cs.LG2025

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Hjalmar Wijk, Tao Lin, Joel Becker +20

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…

cs.AI2025

HCAST: Human-Calibrated Autonomy Software Tasks

David Rein, Joel Becker, Amy Deng +19

To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…