most citedMeasuring AI Ability to Complete Long Software Tasks

14 citations · 14 across the 2 of their papers we have counts for

collaborators

6 papers

cs.AI202614 cited

Measuring AI Ability to Complete Long Software Tasks

Thomas Kwa, Ben West, Joel Becker +23

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…

cs.CY2026

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Stephen Casper, Oskar Galeev, Yoshua Bengio +117

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…

cs.AI2025

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio, Tegan Maharaj, Luke Ong +84

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…

cs.LG2025

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Hjalmar Wijk, Tao Lin, Joel Becker +20

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations fo…

cs.AI2025

HCAST: Human-Calibrated Autonomy Software Tasks

David Rein, Joel Becker, Amy Deng +19

To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…

cs.CL2025

DarkBench: Benchmarking Dark Patterns in Large Language Models

Esben Kran, Hieu Minh "Jord" Nguyen, Akash Kundu +3

We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language mode…