10 citations · 10 across the 4 of their papers we have counts for
1 paper · 1 filter
Zehua Cheng, Jianwei Yang, Wei Dai +1
Large Language Models (LLMs) remain vulnerable to adaptive jailbreaks that easily bypass empirical defenses like GCG. We propose a framework for certifiable robustness that shifts…