4 papers
Code Monitor Red Teaming for Public-Test-Passing Code
Junchi Liao, Jiawen Deng, Fuji Ren
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-like monitoring problem: after code has p…
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Monika JotautaitÄ, Maria Angelica Martinez, Ollie Matthews +1
We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate moni…
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1
As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introd…
Honesty to Subterfuge: In-Context Reinforcement Learning Can Make Honest Models Reward Hack
Leo McKee-Reid, Christoph Sträter, Maria Angelica Martinez +2
Previous work has shown that training "helpful-only" LLMs with reinforcement learning on a curriculum of gameable environments can lead models to generalize to egregious specificat…