most citedMeasuring AI Ability to Complete Long Software Tasks

14 citations · 14 across the 2 of their papers we have counts for

collaborators

5 papers

cs.AI202614 cited

Measuring AI Ability to Complete Long Software Tasks

Thomas Kwa, Ben West, Joel Becker +23

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…

cs.AI2026

MirrorCode: AI can rebuild entire programs from behavior alone

Tom Adamczewski, David Owen, David Rein +4

AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding bench…

cs.CL2025

Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation

Song Wang, Zihan Chen, Peng Wang +5

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or special…

cs.AI2025

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

Joel Becker, Nate Rush, Elizabeth Barnes +1

Despite widespread adoption, the impact of AI tools on software development in the wild remains understudied. We conduct a randomized controlled trial (RCT) to understand how AI to…

cs.AI2025

HCAST: Human-Calibrated Autonomy Software Tasks

David Rein, Joel Becker, Amy Deng +19

To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…