1 paper · 1 filter
Chuhan Zhang, Ye Zhang, Bowen Shi +5
Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechan…