most citedMeasurement to Meaning: A Validity-Centered Framework for AI Evaluation

4 citations · 4 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CY2026

The Limits of AI Data Transparency Policy: Three Disclosure Fallacies

Judy Hanwen Shen, Ken Liu, Angelina Wang +7

Data transparency has emerged as a rallying cry for addressing concerns about AI: data quality, privacy, and copyright chief among them. Yet while these calls are crucial for accou…

cs.CY2025

Disclosure and Evaluation as Fairness Interventions for General-Purpose AI

Vyoma Raman, Judy Hanwen Shen, Andy K. Zhang +4

Despite conflicting definitions and conceptions of fairness, AI fairness researchers broadly agree that fairness is context-specific. However, when faced with general-purpose AI, w…

cs.CL2025

The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior

Angelina Wang, Daniel E. Ho, Sanmi Koyejo

Standard offline evaluations for language models -- a series of independent, state-less inferences made by models -- fail to capture how language models actually behave in practice…

cs.CY20254 cited

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Olawale Salaudeen, Anka Reuel, Ahmed Ahmed +6

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning ca…

cs.CY2025

The California Report on Frontier AI Policy

Rishi Bommasani, Scott R. Singer, Ruth E. Appel +20

The innovations emerging at the frontier of artificial intelligence (AI) are poised to create historic opportunities for humanity but also raise complex policy challenges. Continue…

cs.AI2025

Toward an Evaluation Science for Generative AI Systems

Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach +7

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluatio…