Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework
David Huang, Jaewon Chang, Avidan Shah +2
The Rapid Response (RR) framework, deployed in production systems, including Anthropic's ASL-3 safeguards, continuously improves jailbreak-detection classifiers. When new jailbreak…
cs.LG2023
PubDef: Defending Against Transfer Attacks From Public Models
Chawin Sitawarin, Jaewon Chang, David Huang +2
Adversarial attacks have been a looming and unaddressed threat in the industry. However, through a decade-long history of the robustness evaluation literature, we have learned that…