1 paper · 1 filter
Bang An, Yibo Yang, Dandan Guo +3
Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign ta…