Publications (6)
Taken out of context: On measuring situational awareness in LLMs
Lukas Berglund, Asa Cooper Stickland, Mikita Balesni +5
We aim to better understand the emergence of `situational awareness' in large language models (LLMs). A model is situationally aware if it's aware that it's a model and can recogni…
Visibility into AI Agents
Alan Chan, Carson Ezell, Max Kaufmann +9
Increased delegation of commercial, scientific, governmental, and personal activities to AI agents -- systems capable of pursuing complex goals with limited supervision -- may exac…
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
Max Kaufmann, David Lindner, Roland S. Zimmermann +1
Chain-of-Thought (CoT) monitoring, in which automated systems monitor the CoT of an LLM, is a promising approach for effectively overseeing AI systems. However, the extent to which…
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
Lukas Berglund, Meg Tong, Max Kaufmann +4
We expose a surprising failure of generalization in auto-regressive large language models (LLMs). If a model is trained on a sentence of the form "A is B", it will not automaticall…
Self-Regulation and Requesting Interventions
So Yeon Min, Yue Wu, Jimin Sun +4
Human intelligence involves metacognitive abilities like self-regulation, recognizing limitations, and seeking assistance only when needed. While LLM Agents excel in many domains,…
Testing Robustness Against Unforeseen Adversaries
Max Kaufmann, Daniel Kang, Yi Sun +9
Adversarial robustness research primarily focuses on L_p perturbations, and most defenses are developed with identical training-time and test-time adversaries. However, in real-wor…