activity
20242026
collaborators

6 papers

cs.AI2026

From AGI to ASI

Tim Genewein, Matija Franklin, Alexander Lerchner +11

Over the last decade, building human-level artificial general intelligence has moved from far-fetched speculation to being a concrete next-decade target for many of the largest AI…

cs.AI2026

Control Tax: The Price of Keeping AI in Check

Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre +1

The rapid integration of agentic AI into high-stakes real-world applications requires robust oversight mechanisms. The emerging field of AI Control (AIC) aims to provide such an ov…

cs.AI2025

A Rosetta Stone for AI Benchmarks

Anson Ho, Jean-Stanislas Denain, David Atanasov +2

Most AI benchmarks saturate within years or even months after they are introduced, making it hard to study long-run trends in AI capabilities. To address this challenge, we build a…

cs.CL2025

Inverse Constitutional AI: Compressing Preferences into Principles

Arduin Findeis, Timo Kaufmann, Eyke Hüllermeier +2

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options,…

cs.AI2025

An Approach to Technical AGI Safety and Security

Rohin Shah, Alex Irpan, Alexander Matt Turner +27

Artificial General Intelligence (AGI) promises transformative benefits but also presents significant risks. We develop an approach to address the risk of harms consequential enough…

cs.LG2024

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton, Noah Y. Siegel, János Kramár +8

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…