1 paper
Da Cheng Gu, Yifei Dong, Xinghao Yang +2
Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the…