1 paper · 1 filter
Zhanhao Hu, Julien Piet, Geng Zhao +2
Current LLMs are generally aligned to follow safety requirements and tend to refuse toxic prompts. However, LLMs can fail to refuse toxic prompts or be overcautious and refuse beni…