1 paper · 1 filter
Sha Luo, Sang Jung Kim, Zening Duan +1
Refusal behavior by Large Language Models is increasingly visible in content moderation, yet little is known about how refusals vary by the identity of the user making the request.…