6 papers · 1 filter
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Jack Hopkins, Dipika Khullar, Fabien Roger
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information…
Classifier Context Rot: Monitor Performance Degrades with Context Length
Sam Martin, Fabien Roger
Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely c…
How Useful Is Cross-Domain Generalization for Training LLM Monitors?
Sam Martin, Fabien Roger
Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tun…
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
Dipika Khullar, Jack Hopkins, Rowan Wang +1
Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess…
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright +17
We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…