activity
20242026
collaborators

6 papers

cs.CL2026

ASSERT: A Measurement Pipeline for GenAI Audits

Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11

Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…

cs.LG2026

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

Alexandra Chouldechova, A. Feder Cooper, Solon Barocas +3

We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR)…

cs.CL2025

Anecdoctoring: Automated Red-Teaming Across Language and Place

Alejandro Cuevas, Saloni Dash, Bharat Kumar Nayak +2

Disinformation is among the top risks of generative artificial intelligence (AI) misuse. Global adoption of generative AI necessitates red-teaming evaluations (i.e., systematic adv…

cs.CL2025

Taxonomizing Representational Harms using Speech Act Theory

Emily Corvi, Hannah Washington, Stefanie Reed +9

Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a…

cs.CY2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper +17

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…

cs.CY2024

A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

Alexandra Chouldechova, Chad Atalla, Solon Barocas +11

The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard…