2 citations · 6 across the 9 of their papers we have counts for
11 papers
Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming
Alexandra Chouldechova, A. Feder Cooper, Solon Barocas +3
We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR)…
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems
Emma Harvey, Emily Sheng, Su Lin Blodgett +4
The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…
Taxonomizing Representational Harms using Speech Act Theory
Emily Corvi, Hannah Washington, Stefanie Reed +9
Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a…
Validating LLM-as-a-Judge Systems under Rating Indeterminacy
Luke Guerdan, Solon Barocas, Kenneth Holstein +3
The LLM-as-a-judge paradigm, in which a judge LLM system replaces human raters in rating the outputs of other generative AI (GenAI) systems, plays a critical role in scaling and st…
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Hanna Wallach, Meera Desai, A. Feder Cooper +17
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…
A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts
Alexandra Chouldechova, Chad Atalla, Solon Barocas +11
The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard…