9 papers
Evading Chain-of-Thought Monitoring Through Model Poisoning
Giorgio Severi, Shujaat Mirza, Blake Bullwinkel +1
Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its ac…
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…
Beyond the Single Turn: Reframing Refusals as Dynamic Experiences Embedded in the Context of Mental Health Support Interactions with LLMs
Ningjing Tang, Alice Qian, Qiaosi Wang +6
Content Warning: This paper contains participant quotes and discussions related to mental health challenges, emotional distress, and suicidal ideation. Large language models (LLMs)…
GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Mark Russinovich, Yanan Cai, Keegan Hines +3
Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-…
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
Blake Bullwinkel, Giorgio Severi, Keegan Hines +3
Detecting whether a model has been poisoned is a longstanding problem in AI security. In this work, we present a practical scanner for identifying sleeper agent-style backdoors in…
A Systematization of Security Vulnerabilities in Computer Use Agents
Daniel Jones, Giorgio Severi, Martin Pouliot +7
Computer Use Agents (CUAs), autonomous systems that interact with software interfaces via browsers or virtual machines, are rapidly being deployed in consumer and enterprise enviro…