1 paper
Sha Luo, Sang Jung Kim, Zening Duan +1
Refusal behavior by Large Language Models is increasingly visible in content moderation, yet little is known about how refusals vary by the identity of the user making the request.…