7 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CR2026
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
Hoagy Cunningham, Jerry Wei, Zihan Wang +26
We introduce enhanced Constitutional Classifiers that deliver production-grade jailbreak robustness with dramatically reduced computational costs and refusal rates compared to prev…
cs.CL2025★ 7 cited
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma, Meg Tong, Jesse Mu +40
Large language models (LLMs) are vulnerable to universal jailbreaks-prompting strategies that systematically bypass model safeguards and enable users to carry out harmful processes…
stat.ML2021★ 2 cited
Robust Semantic Interpretability: Revisiting Concept Activation Vectors
Jacob Pfau, Albert T. Young, Jerome Wei +2
Interpretability methods for image classification assess model trustworthiness by attempting to expose whether the model is systematically biased or attending to the same cues as a…