1 paper
Utsav Maskey, Sumit Yadav, Mark Dras +1
LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful content. This phenomenon diminis…