Publications (6)
Evaluating Language-Model Agents on Realistic Autonomous Tasks
Megan Kinniment, Lucas Jun Koba Sato, Haoxing Du +10
In this report, we explore the ability of language model agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. We refe…
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker +23
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…
The science and practice of proportionality in AI risk evaluations
Carlos Mougan, Lauritz Morlock, Jair Aguirre +19
A global challenge in artificial intelligence (AI) regulation lies in achieving effective risk management without compromising innovation and technical progress. The European Union…
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
Ahmad Idrissi-Yaghir, Amin Dada, Henning Schäfer +17
Recent advances in natural language processing (NLP) can be largely attributed to the advent of pre-trained language models such as BERT and RoBERTa. While these models demonstrate…
AutoPET Challenge: Combining nn-Unet with Swin UNETR Augmented by Maximum Intensity Projection Classifier
Lars Heiliger, Zdravko Marinov, Max Hasin +9
Tumor volume and changes in tumor characteristics over time are important biomarkers for cancer therapy. In this context, FDG-PET/CT scans are routinely used for staging and re-sta…
HCAST: Human-Calibrated Autonomy Software Tasks
David Rein, Joel Becker, Amy Deng +19
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…