1 paper · 1 filter
Francisco Eiras, Aleksandar Petrov, Philip H. S. Torr +2
Recent research shows that fine-tuning on benign instruction-following data can inadvertently undo the safety alignment process and increase a model's propensity to comply with har…