6 papers
Diffuse AI Control on Fuzzy Tasks
Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar +1
AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned wit…
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
Jinghan Jia, Joe Benton, Eric Easley
Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produces the final answer. However,…
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Elle Najt, Colin Toft, Tyler Tracy +2
Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To…
Removing Sandbagging in LLMs by Training with Weak Supervision
Emil Ryd, Henning Bartsch, Julian Stastny +2
As AI systems begin to automate complex tasks, supervision increasingly relies on weaker models or limited human oversight that cannot fully verify output quality. A model more cap…
Evaluating Control Protocols for Untrusted AI Agents
Jon Kutasov, Chloe Loughridge, Yuqi Sun +4
As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untruste…
Optimizing AI Agent Attacks With Synthetic Data
Chloe Loughridge, Paul Colognese, Avery Griffin +3
As AI deployments become more complex and high-stakes, it becomes increasingly important to be able to estimate their risk. AI control is one framework for doing so. However, good…