1 paper · 1 filter
Tianle Gu, Kexin Huang, Zongqi Wang +7
Safety alignment is a key requirement for building reliable Artificial General Intelligence. Despite significant advances in safety alignment, we observe that minor latent shifts c…