1 paper · 1 filter
Junlong Liu, Haobo Wang, Weiqi Luo +1
Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and defenses at the prompt level,…