6 citations · 17 across the 12 of their papers we have counts for
3 papers · 1 filter
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Blake Bullwinkel, Eugenia Kim, Amanda Minnich +1
AI red teaming must continually adapt to evolving attackers and defenders. Reinforcement learning offers a promising approach to discovering novel attacks, and co-training methods…
Jailbreak Distillation: Renewable Safety Benchmarking
Jingyu Zhang, Ahmed Elgohary, Xiawei Wang +5
Large language models (LLMs) are rapidly deployed in critical applications, raising urgent needs for robust safety benchmarking. We propose Jailbreak Distillation (JBDistill), a no…
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
Emman Haider, Daniel Perez-Becker, Thomas Portet +28
Recent innovations in language model training have demonstrated that it is possible to create highly performant models that are small enough to run on a smartphone. As these models…