19 citations · 31 across the 11 of their papers we have counts for
1 paper · 1 filter
Zhehao Zhang, Weijie Xu, Fanyou Wu +1
Safety alignment approaches in large language models (LLMs) often lead to the over-refusal of benign queries, significantly diminishing their utility in sensitive scenarios. To add…