1 paper
Bocheng Chen, Xiaoqun Liu, Hanqing Guo +2
Large language models (LLMs) remain vulnerable to jailbreak attacks in which adversarial prompts induce harmful outputs. Existing defenses often require access to the model interna…