activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget

Florian E. Dorner, Moritz Hardt

We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a giv…

cs.LG2025

ImageNot: A contrast with ImageNet preserves model rankings

Olawale Salaudeen, Moritz Hardt

We introduce ImageNot, a dataset constructed explicitly to be drastically different than ImageNet while matching its scale. ImageNot is designed to test the external validity of de…

cs.LG2025

Performative Prediction: Past and Future

Moritz Hardt, Celestine Mendler-Dünner

Predictions in the social world generally influence the target of prediction, a phenomenon known as performativity. Self-fulfilling and self-negating predictions are examples of pe…

cs.LG2024

What Makes ImageNet Look Unlike LAION

Ali Shirali, Moritz Hardt

ImageNet was famously created from Flickr image search results. What if we recreated ImageNet instead by searching the massive LAION dataset based on image captions alone? In this…

cs.LG2024

Do causal predictors generalize better to new domains?

Vivian Y. Nastl, Moritz Hardt

We study how well machine learning models trained on causal features generalize across domains. We consider 16 prediction tasks on tabular datasets covering applications in health,…

cs.LG2024

Evaluating language models as risk scores

André F. Cruz, Moritz Hardt, Celestine Mendler-Dünner

Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the…