2 papers
cs.CL2026
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic +5
Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate these choices using two…
cs.CL2026
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models
Mengya Hu, Qiong Wei, Sandeep Atluri
Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide…