8 papers
Evading Chain-of-Thought Monitoring Through Model Poisoning
Giorgio Severi, Shujaat Mirza, Blake Bullwinkel +1
Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning trace is informative about its ac…
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Dasol Choi, Eugenia Kim, Jaewon Noh +14
Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's a…
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
Blake Bullwinkel, Giorgio Severi, Keegan Hines +3
Detecting whether a model has been poisoned is a longstanding problem in AI security. In this work, we present a practical scanner for identifying sleeper agent-style backdoors in…
A Systematization of Security Vulnerabilities in Computer Use Agents
Daniel Jones, Giorgio Severi, Martin Pouliot +7
Computer Use Agents (CUAs), autonomous systems that interact with software interfaces via browsers or virtual machines, are rapidly being deployed in consumer and enterprise enviro…
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
Blake Bullwinkel, Mark Russinovich, Ahmed Salem +8
Recent research has demonstrated that state-of-the-art LLMs and defenses remain susceptible to multi-turn jailbreak attacks. These attacks require only closed-box model access and…