6 citations · 6 across the 2 of their papers we have counts for
1 paper · 1 filter
Yu Feng, Chunting Zang, Chen Shen +4
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard…