7 papers
How Transparent is DiffusionGemma?
Joshua Engels, Callum McDougall, Bilal Chughtai +13
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, Diffus…
Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins, Bilal Chughtai, Joshua Engels
Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their…
Building Production-Ready Probes For Gemini
János Kramár, Joshua Engels, Zheng Wang +4
Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…
Difficulties with Evaluating a Deception Detector for AIs
Lewis Smith, Bilal Chughtai, Neel Nanda
Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evid…
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Zora Che, Stephen Casper, Robert Kirk +12
Evaluations of large language model (LLM) risks and capabilities are increasingly being incorporated into AI risk management and governance frameworks. Currently, most risk evaluat…
Detecting Strategic Deception Using Linear Probes
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim +1
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs…