1 paper
Riad Ahmed Anonto, Md Labid Al Nahiyan, Md Tanvir Hassan
Safety-aligned language models often refuse prompts that are actually harmless. Current evaluations mostly report global rates such as false rejection or compliance. These scores t…