4 papers
Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework
David Huang, Jaewon Chang, Avidan Shah +2
The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers. When new jailbreak…
Covert Influence Between Language Models
Avidan Shah, Jay Chooi, Jinghua Ou +1
As language models increasingly consume one another's outputs, covert influence -- a phenomenon where a sender's payload (the behavioral disposition it is conditioned to propagate)…
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
Avidan Shah, Jannik Brinkmann, Rico Angell
As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection…
On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation
Andy Han, Kristina Fujimoto, Avidan Shah +5
Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promi…