1 paper
Seunghyun Lee, Dongyoon Han, Sangdoo Yun
Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal…