1 paper · 1 filter
Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai +2
Production LLMs increasingly rely on toxicity-based moderation filters as a primary defense, assuming that harmful intent correlates with toxic surface wording. We show this assump…