2 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Qinyan Zhou, Peixin Zhang, Jun Sun +2
While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that…