most citedMeasurement to Meaning: A Validity-Centered Framework for AI Evaluation

4 citations · 4 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG2025

Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations

Olawale Salaudeen, Haoran Zhang, Kumail Alhamoud +2

Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-…

cs.AI2025

A Definition of AGI

Dan Hendrycks, Dawn Song, Christian Szegedy +30

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…

cs.CY20254 cited

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Olawale Salaudeen, Anka Reuel, Ahmed Ahmed +6

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning ca…

stat.ML2025

Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness

Stephen R. Pfohl, Natalie Harris, Chirag Nagpal +12

Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal pe…

cs.LG2025

Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?

Olawale Salaudeen, Nicole Chiou, Shiny Weng +1

Spurious correlations, unstable statistical shortcuts a model can exploit, are expected to degrade performance out-of-distribution (OOD). However, across many popular OOD generaliz…

cs.AI2025

Toward an Evaluation Science for Generative AI Systems

Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach +7

There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluatio…