From the 2 of 4 linked papers with an AI index.
4 papers
ASSERT: A Measurement Pipeline for GenAI Audits
Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic +5
The paper studies how to best place and combine content‑moderation filters and response rewriting in conversational systems, measuring overall usefulness and harmful exposure rathe…
From Prompt Risk to Response Risk: Paired Analysis of Safety Behavior of Large Language Models
Mengya Hu, Qiong Wei, Sandeep Atluri
The paper proposes a paired analysis framework that compares the risk level of prompts and their LLM-generated responses across multiple harm categories and severity levels, reveal…
AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models
Tanmay Gautam, Alireza Bahramali, Sandeep Atluri
Automated red-teaming methods for large language models typically optimize attack prompts within a fixed, human-designed strategy, leaving the attack strategy itself unchanged. We…