1 paper
Shei Pern Chua, Hao Wu, Qianli Ma +1
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of rob…