4 papers · 1 filter
ASSERT: A Measurement Pipeline for GenAI Audits
Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic +5
Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate these choices using two…
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models
Mengya Hu, Qiong Wei, Sandeep Atluri
Safety evaluations of large language models (LLMs) typically report binary outcomes, i.e. attack success rate (ASR), refusal rate, or harmful versus safe classification, which hide…
Generating Self-Contained and Summary-Centric Question Answer Pairs via Differentiable Reward Imitation Learning
Li Zhou, Kevin Small, Yong Zhang +1
Motivated by suggested question generation in conversational news recommendation systems, we propose a model for generating question-answer pairs (QA pairs) with self-contained, su…