1 paper · 1 filter
Songze Li, Ruishi He, Xiaojun Jia +2
Large Language Models (LLMs) face a significant threat from multi-turn jailbreak attacks, where adversaries progressively steer conversations to elicit harmful outputs. However, th…