papers

Publications (10)

cs.CY2026

The science and practice of proportionality in AI risk evaluations

Carlos Mougan, Lauritz Morlock, Jair Aguirre +19

A global challenge in artificial intelligence (AI) regulation lies in achieving effective risk management without compromising innovation and technical progress. The European Union…

cs.CL2023

How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions

Lorenzo Pacchiardi, Alex J. Chan, Sören Mindermann +5

Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when inst…

cs.CL2023

Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Ansh Radhakrishnan, Karina Nguyen, Anna Chen +21

As large language models (LLMs) perform more difficult tasks, it becomes harder to verify the correctness and safety of their behavior. One approach to help with this issue is to p…

cs.CY2025

Thousands of AI Authors on the Future of AI

Katja Grace, Harlan Stewart, Julia Fabienne Sandkühler +4

In the largest survey of its kind, 2,778 researchers who had published in top-tier artificial intelligence (AI) venues gave predictions on the pace of AI progress and the nature an…

cs.AI2023

Measuring Faithfulness in Chain-of-Thought Reasoning

Tamera Lanham, Anna Chen, Ansh Radhakrishnan +27

Large language models (LLMs) perform better when they produce step-by-step, "Chain-of-Thought" (CoT) reasoning before answering a question, but it is unclear if the stated reasonin…

cs.LG2022

Prioritized Training on Points that are Learnable, Worth Learning, and Not Yet Learnt

Sören Mindermann, Jan Brauner, Muhammed Razzak +8

Training on web-scale data can take months. But most computation and time is wasted on redundant and noisy points that are already learnt or not learnable. To accelerate training,…

cs.CR2024

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Evan Hubinger, Carson Denison, Jesse Mu +36

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when giv…

cs.LG2023

Prioritized training on points that are learnable, worth learning, and not yet learned (workshop version)

Sören Mindermann, Muhammed Razzak, Winnie Xu +7

We introduce Goldilocks Selection, a technique for faster model training which selects a sequence of training points that are "just right". We propose an information-theoretic acqu…

cs.CY2024

Managing extreme AI risks amid rapid progress

Yoshua Bengio, Geoffrey Hinton, Andrew Yao +22

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increase…

cs.AI2022

Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec +2

Benchmarks are crucial to measuring and steering progress in artificial intelligence (AI). However, recent studies raised concerns over the state of AI benchmarking, reporting issu…