9 papers
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
Tom Sühr, Florian E. Dorner, Olawale Salaudeen +2
Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intel…
ImageNot: A contrast with ImageNet preserves model rankings
Olawale Salaudeen, Moritz Hardt
We introduce ImageNot, a dataset constructed explicitly to be drastically different than ImageNet while matching its scale. ImageNot is designed to test the external validity of de…
A Definition of AGI
Dan Hendrycks, Dawn Song, Christian Szegedy +30
The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…
Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Olawale Salaudeen, Haoran Zhang, Kumail Alhamoud +2
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-…
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
Stephen R. Pfohl, Natalie Harris, Chirag Nagpal +12
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal pe…
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
Olawale Salaudeen, Nicole Chiou, Shiny Weng +1
Spurious correlations, unstable statistical shortcuts a model can exploit, are expected to degrade performance out-of-distribution (OOD). However, across many popular OOD generaliz…