1 paper · 1 filter
Somnath Banerjee, Sayan Layek, Soham Tripathy +3
Safety-aligned language models often exhibit fragile and imbalanced safety mechanisms, increasing the likelihood of generating unsafe content. In addition, incorporating new knowle…