1 paper · 1 filter
Rui Wu, Yihao Quan, Zeru Shi +3
Safety-aligned Large Language Models (LLMs) still show two dominant failure modes: they are easily jailbroken, or they over-refuse harmless inputs that contain sensitive surface si…