4 papers · 1 filter
Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework
David Huang, Jaewon Chang, Avidan Shah +2
The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers. When new jailbreak…
Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training
Avidan Shah, Jannik Brinkmann, Rico Angell
As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection…
On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation
Andy Han, Kristina Fujimoto, Avidan Shah +5
Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promi…
Efficient Mitigation of Bus Bunching through Setter-Based Curriculum Learning
Avidan Shah, Danny Tran, Yuhan Tang
Curriculum learning has been growing in the domain of reinforcement learning as a method of improving training efficiency for various tasks. It involves modifying the difficulty (l…