activity
20242026
collaborators

9 papers

cs.CL2026

ASSERT: A Measurement Pipeline for GenAI Audits

Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11

Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…

cs.CL2026

AI-Assisted Systematization for Evaluating GenAI Systems

Dhruv Agarwal, Emily Sheng, Chad Atalla +6

Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as "reasoning," "fairness," or "creativity." When the…

cs.LG2026

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

Alexandra Chouldechova, A. Feder Cooper, Solon Barocas +3

We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR)…

cs.LG2025

Validating LLM-as-a-Judge Systems under Rating Indeterminacy

Luke Guerdan, Solon Barocas, Kenneth Holstein +3

The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and st…

cs.CL2025

Taxonomizing Representational Harms using Speech Act Theory

Emily Corvi, Hannah Washington, Stefanie Reed +9

Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a…

cs.CY2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

Emma Harvey, Emily Sheng, Su Lin Blodgett +4

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…