most citedEvaluating Generative AI Systems is a Social Science Measurement Challenge

2 citations · 4 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CY2025

Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

Emma Harvey, Emily Sheng, Su Lin Blodgett +4

The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instrumen…

cs.CL20251 cited

Taxonomizing Representational Harms using Speech Act Theory

Emily Corvi, Hannah Washington, Stefanie Reed +9

Representational harms are widely recognized among fairness-related harms caused by generative language systems. However, their definitions are commonly under-specified. We make a…

cs.CY2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper +17

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…

cs.CY2024

A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

Alexandra Chouldechova, Chad Atalla, Solon Barocas +11

The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard…

cs.CY20241 cited

Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems

Emma Harvey, Emily Sheng, Su Lin Blodgett +4

To facilitate the measurement of representational harms caused by large language model (LLM)-based systems, the NLP research community has produced and made publicly available nume…

cs.CY20242 cited

Evaluating Generative AI Systems is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, Nicholas Pangakis +17

Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult…