activity
20242026
most citedEvaluating Generative AI Systems is a Social Science Measurement Challenge

2 citations · 5 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Comparison requires valid measurement: Rethinking attack success rate comparisons in AI red teaming

Alexandra Chouldechova, A. Feder Cooper, Solon Barocas +3

We argue that conclusions drawn about relative system safety or attack method efficacy via AI red teaming are often not supported by evidence provided by attack success rate (ASR)…

cs.CL20262 cited

Extracting books from production language models

Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo +1

Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether…

cs.HC20251 cited

XR Blocks: Accelerating Human-centered AI + XR Innovation

David Li, Nels Numan, Xun Qian +17

We are on the cusp where Artificial Intelligence (AI) and Extended Reality (XR) are converging to unlock new paradigms of interactive computing. However, a significant gap exists b…

cs.CY2025

The California Report on Frontier AI Policy

Rishi Bommasani, Scott R. Singer, Ruth E. Appel +20

The innovations emerging at the frontier of artificial intelligence (AI) are poised to create historic opportunities for humanity but also raise complex policy challenges. Continue…

cs.CY2025

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Hanna Wallach, Meera Desai, A. Feder Cooper +17

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] a…

cs.CY2024

A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts

Alexandra Chouldechova, Chad Atalla, Solon Barocas +11

The valid measurement of generative AI (GenAI) systems' capabilities, risks, and impacts forms the bedrock of our ability to evaluate these systems. We introduce a shared standard…