activity
20242026
collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

Jack Hopkins, Dipika Khullar, Fabien Roger

Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information…

cs.AI2026

Classifier Context Rot: Monitor Performance Degrades with Context Length

Sam Martin, Fabien Roger

Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely c…

cs.AI2026

How Useful Is Cross-Domain Generalization for Training LLM Monitors?

Sam Martin, Fabien Roger

Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tun…

cs.AI2026

Self-Attribution Bias: When AI Monitors Go Easy on Themselves

Dipika Khullar, Jack Hopkins, Rowan Wang +1

Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess…

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI2024

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…