1 paper · 1 filter
Jinman Wu, Yi Xie, Shen Lin +2
Safety alignment is often conceptualized as a monolithic process wherein harmfulness detection automatically triggers refusal. However, the persistence of jailbreak attacks suggest…