8 papers
ASSERT: A Measurement Pipeline for GenAI Audits
Riccardo Fogliato, Abhinav Palia, Xiawei Wang +11
Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate t…
AI-Assisted Systematization for Evaluating GenAI Systems
Dhruv Agarwal, Emily Sheng, Chad Atalla +6
Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as "reasoning," "fairness," or "creativity." When the…
Taxonomizing Representational Harms using Speech Act Theory
Emily Corvi, Hannah Washington, Stefanie Reed +9
Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a…
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper +17
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
Emma Harvey, Emily Sheng, Su Lin Blodgett +4
The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
Alexandra Chouldechova, Chad Atalla, Solon Barocas +11
The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard…