Publications (91)
Alternative Microfoundations for Strategic Classification
Meena Jagadeesan, Celestine Mendler-Dünner, Moritz Hardt
When reasoning about strategic behavior in a machine learning context it is tempting to combine standard microfoundations of rational agents with the statistical decision theory un…
From Optimizing Engagement to Measuring Value
Smitha Milli, Luca Belli, Moritz Hardt
Most recommendation engines today are based on predicting user engagement, e.g. predicting whether a user will click on an item or not. However, there is potentially a large gap be…
Model Reconstruction from Model Explanations
Smitha Milli, Ludwig Schmidt, Anca D. Dragan +1
We show through theory and experiment that gradient-based explanations of a model quickly reveal the model itself. Our results speak to a tension between the desire to keep a propr…
Generalization in Adaptive Data Analysis and Holdout Reuse
Cynthia Dwork, Vitaly Feldman, Moritz Hardt +3
Overfitting is the bane of data analysts, even when data are plentiful. Formal approaches to understanding this problem focus on statistical inference and generalization of individ…
Difficult Lessons on Social Prediction from Wisconsin Public Schools
Juan C. Perdomo, Tolani Britton, Moritz Hardt +1
Early warning systems (EWS) are predictive tools at the center of recent efforts to improve graduation rates in public schools across the United States. These systems assist in tar…
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
Multiple choice benchmarks have long been the workhorse of language model evaluation because grading multiple choice is objective and easy to automate. However, we show multiple ch…
Strategic Classification is Causal Modeling in Disguise
John Miller, Smitha Milli, Moritz Hardt
Consequential decision-making incentivizes individuals to strategically adapt their behavior to the specifics of the decision rule. While a long line of work has viewed strategic a…
Leaderboard Incentives: Model Rankings under Strategic Post-Training
Yatong Chen, Guanhua Zhang, Moritz Hardt
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmax…
Training on the Test Task Confounds Evaluation and Emergence
Ricardo Dominguez-Olmedo, Florian E. Dorner, Moritz Hardt
We study a fundamental problem in the evaluation of large language models that we call training on the test task. Unlike wrongful practices like training on the test data, leakage,…
Stable Recurrent Models
John Miller, Moritz Hardt
Stability is a fundamental property of dynamical systems, yet to this date it has had little bearing on the practice of recurrent neural networks. In this work, we conduct a thorou…
Algorithmic Collective Action in Machine Learning
Moritz Hardt, Eric Mazumdar, Celestine Mendler-Dünner +1
We initiate a principled study of algorithmic collective action on digital platforms that deploy machine learning algorithms. We propose a simple theoretical model of a collective…
The advantages of multiple classes for reducing overfitting from test set reuse
Vitaly Feldman, Roy Frostig, Moritz Hardt
Excessive reuse of holdout data can lead to overfitting. However, there is little concrete evidence of significant overfitting due to holdout reuse in popular multiclass benchmarks…
Unprocessing Seven Years of Algorithmic Fairness
André F. Cruz, Moritz Hardt
Seven years ago, researchers proposed a postprocessing method to equalize the error rates of a model across different demographic groups. The work launched hundreds of papers purpo…
Preserving Statistical Validity in Adaptive Data Analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt +3
A great deal of effort has been devoted to reducing the risk of spurious scientific discoveries, from the use of sophisticated validation techniques, to deep statistical methods fo…
Computational Arbitrage in AI Model Markets
Ricardo Olmedo, Bernhard Schölkopf, Moritz Hardt
Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a…
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
Jonas Hübotter, Leander Diaz-Bone, Ido Hakimi +2
Humans are good at learning on the job: We learn how to solve the tasks we face as we go along. Can a model do the same? We propose an agent that assembles a task-specific curricul…
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt +2
Despite their massive size, successful deep artificial neural networks can exhibit a remarkably small difference between training and test performance. Conventional wisdom attribut…
ImageNot: A contrast with ImageNet preserves model rankings
Olawale Salaudeen, Moritz Hardt
We introduce ImageNot, a dataset constructed explicitly to be drastically different than ImageNet while matching its scale. ImageNot is designed to test the external validity of de…
How Robust are Linear Sketches to Adaptive Inputs?
Moritz Hardt, David P. Woodruff
Linear sketches are powerful algorithmic tools that turn an n-dimensional input into a concise lower-dimensional representation via a linear transformation. Such sketches have seen…
Retraining Seeks Stable Signals
Moritz Hardt
Predictive models deployed at scale influence future data, a phenomenon called performativity. And there is always one way to cope: Train the model on new data, deploy it again, an…
Gradient Descent Learns Linear Dynamical Systems
Moritz Hardt, Tengyu Ma, Benjamin Recht
We prove that stochastic gradient descent efficiently converges to the global optimizer of the maximum likelihood objective of an unknown linear time-invariant dynamical system fro…
Do causal predictors generalize better to new domains?
Vivian Y. Nastl, Moritz Hardt
We study how well machine learning models trained on causal features generalize across domains. We consider 16 prediction tasks on tabular datasets covering applications in health,…
Decline Now: A Combinatorial Model for Algorithmic Collective Action
Dorothee Sigg, Moritz Hardt, Celestine Mendler-Dünner
Drivers on food delivery platforms often run a loss on low-paying orders. In response, workers on DoorDash started a campaign, #DeclineNow, to purposefully decline orders below a c…
Tight bounds for learning a mixture of two gaussians
Moritz Hardt, Eric Price
We consider the problem of identifying the parameters of an unknown mixture of two arbitrary -dimensional gaussians from a sequence of independent random samples. Our main resul…
Beyond Worst-Case Analysis in Private Singular Vector Computation
Moritz Hardt, Aaron Roth
We consider differentially private approximate singular vector computation. Known worst-case lower bounds show that the error of any differentially private algorithm must scale pol…
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, Yoram Singer
We show that parametric models trained by a stochastic gradient method (SGM) with few iterations have vanishing generalization error. We prove our results by arguing that SGM is al…
Allocation Requires Prediction Only if Inequality Is Low
Ali Shirali, Rediet Abebe, Moritz Hardt
Algorithmic predictions are emerging as a promising solution concept for efficiently allocating societal resources. Fueling their use is an underlying assumption that such systems…
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
Yu Sun, Xiaolong Wang, Zhuang Liu +3
In this paper, we propose Test-Time Training, a general approach for improving the performance of predictive models when training and test data come from different distributions. W…
Algorithms and Hardness for Robust Subspace Recovery
Moritz Hardt, Ankur Moitra
We consider a fundamental problem in unsupervised learning called \emph{subspace recovery}: given a collection of points in , if many but not necessarily all of t…
Avoiding Discrimination through Causal Reasoning
Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo +3
Recent work on fairness in machine learning has focused on various statistical discrimination criteria and how they trade off. Most of these criteria are observational: They depend…
Climbing a shaky ladder: Better adaptive risk estimation
Moritz Hardt
We revisit the \emph{leaderboard problem} introduced by Blum and Hardt (2015) in an effort to reduce overfitting in machine learning benchmarks. We show that a randomized version o…
Questioning the Survey Responses of Large Language Models
Ricardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-Dünner
Surveys have recently gained popularity as a tool to study large language models. By comparing survey responses of models to those of human reference populations, researchers aim t…
Strategic Classification
Moritz Hardt, Nimrod Megiddo, Christos Papadimitriou +1
Machine learning relies on the assumption that unseen test instances of a classification problem follow the same distribution as observed training data. However, this principle can…
Performative Prediction
Juan C. Perdomo, Tijana Zrnic, Celestine Mendler-Dünner +1
When predictions support decisions they may influence the outcome they aim to predict. We call such predictions performative; the prediction influences the target. Performativity i…
Patterns, predictions, and actions: A story about machine learning
Moritz Hardt, Benjamin Recht
This graduate textbook on machine learning tells a story of how patterns in data support predictions and consequential actions. Starting with the foundations of decision making, we…
Preventing False Discovery in Interactive Data Analysis is Hard
Moritz Hardt, Jonathan Ullman
We show that, under a standard hardness assumption, there is no computationally efficient algorithm that given samples from an unknown distribution can give valid answers to $n…
The Social Cost of Strategic Classification
Smitha Milli, John Miller, Anca D. Dragan +1
Consequential decision-making typically incentivizes individuals to behave strategically, tailoring their behavior to the specifics of the decision rule. A long line of work has th…
Subsampling Mathematical Relaxations and Average-case Complexity
Boaz Barak, Moritz Hardt, Thomas Holenstein +1
We initiate a study of when the value of mathematical relaxations such as linear and semidefinite programs for constraint satisfaction problems (CSPs) is approximately preserved wh…
Good Allocations from Bad Estimates
SÃlvia Casacuberta, Moritz Hardt
Conditional average treatment effect (CATE) estimation is the de facto gold standard for targeting a treatment to a heterogeneous population. The method estimates treatment effects…
Lawma: The Power of Specialization for Legal Annotation
Ricardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe +6
Annotation and classification of legal text are central components of empirical legal research. Traditionally, these tasks are often delegated to trained research assistants. Motiv…
The implicit fairness criterion of unconstrained learning
Lydia T. Liu, Max Simchowitz, Moritz Hardt
We clarify what fairness guarantees we can and cannot expect to follow from unconstrained machine learning. Specifically, we characterize when unconstrained learning on its own imp…
Beating Randomized Response on Incoherent Matrices
Moritz Hardt, Aaron Roth
Computing accurate low rank approximations of large matrices is a fundamental data mining task. In many applications however the matrix contains sensitive information about individ…
Train-before-Test Harmonizes Language Model Rankings
Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt
Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model…
Performative Prediction: Past and Future
Moritz Hardt, Celestine Mendler-Dünner
Predictions in the social world generally influence the target of prediction, a phenomenon known as performativity. Self-fulfilling and self-negating predictions are examples of pe…
An engine not a camera: Measuring performative power of online search
Celestine Mendler-Dünner, Gabriele Carovano, Moritz Hardt
The power of digital platforms is at the center of major ongoing policy and regulatory efforts. To advance existing debates, we designed and executed an experiment to measure the p…
A System for Massively Parallel Hyperparameter Tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh +4
Modern learning models are characterized by large hyperparameter spaces and long training times. These properties, coupled with the rise of parallel computing and the growing deman…
Policy Design in Long-Run Welfare Dynamics
Jiduan Wu, Rediet Abebe, Moritz Hardt +1
Improving social welfare is a complex challenge requiring policymakers to optimize objectives across multiple time horizons. Evaluating the impact of such policies presents a funda…
Test-Time Training on Nearest Neighbors for Large Language Models
Moritz Hardt, Yu Sun
Many recent efforts augment language models with retrieval, by adding retrieved data to the input context. For this approach to succeed, the retrieved data must be added at both tr…
A Theory of Dynamic Benchmarks
Ali Shirali, Rediet Abebe, Moritz Hardt
Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks. In contrast to an extensive theoretical and empirica…
Causal Inference from Competing Treatments
Ana-Andreea Stoica, Vivian Y. Nastl, Moritz Hardt
Many applications of RCTs involve the presence of multiple treatment administrators -- from field experiments to online advertising -- that compete for the subjects' attention. In…
Delayed Impact of Fair Machine Learning
Lydia T. Liu, Sarah Dean, Esther Rolf +2
Fairness in machine learning has predominantly been studied in static classification settings without concern for how decisions change the underlying population over time. Conventi…
The Ladder: A Reliable Leaderboard for Machine Learning Competitions
Avrim Blum, Moritz Hardt
The organizer of a machine learning competition faces the problem of maintaining an accurate leaderboard that faithfully represents the quality of the best submission of each compe…
Stochastic wage suppression on gig platforms and how to organize against it
Ana-Andreea Stoica, Celestine Mendler-Duenner, Moritz Hardt
Digital labor platforms are increasingly used to procure human input, ranging from annotating data and red-teaming AI models, to ride-sharing and food delivery. A central concern i…
Understanding Alternating Minimization for Matrix Completion
Moritz Hardt
Alternating Minimization is a widely used and empirically successful heuristic for matrix completion and related low-rank optimization problems. Theoretical guarantees for Alternat…
Causal Inference Struggles with Agency on Online Platforms
Smitha Milli, Luca Belli, Moritz Hardt
Online platforms regularly conduct randomized experiments to understand how changes to the platform causally affect various outcomes of interest. However, experimentation on online…
Revisiting Design Choices in Proximal Policy Optimization
Chloe Ching-Yun Hsu, Celestine Mendler-Dünner, Moritz Hardt
Proximal Policy Optimization (PPO) is a popular deep policy gradient algorithm. In standard implementations, PPO regularizes policy updates with clipped probability ratios, and par…
Algorithmic Amplification of Politics on Twitter
Ferenc Huszár, Sofia Ira Ktena, Conor O'Brien +3
Content on Twitter's home timeline is selected and ordered by personalization algorithms. By consistently ranking certain content higher, these algorithms may amplify some messages…
Sanity Checks for Saliency Maps
Julius Adebayo, Justin Gilmer, Michael Muelly +3
Saliency methods have emerged as a popular tool to highlight features in an input deemed relevant for the prediction of a learned model. Several saliency methods have been proposed…
Natural Analysts in Adaptive Data Analysis
Tijana Zrnic, Moritz Hardt
Adaptive data analysis is frequently criticized for its pessimistic generalization guarantees. The source of these pessimistic bounds is a model that permits arbitrary, possibly ad…
The Noisy Power Method: A Meta Algorithm with Applications
Moritz Hardt, Eric Price
We provide a new robust convergence analysis of the well-known power method for computing the dominant singular vectors of a matrix that we call the noisy power method. Our result…
Adversarial Scrutiny of Evidentiary Statistical Software
Rediet Abebe, Moritz Hardt, Angela Jin +3
The U.S. criminal legal system increasingly relies on software output to convict and incarcerate people. In a large number of cases each year, the government makes these consequent…
Inherent Trade-Offs between Diversity and Stability in Multi-Task Benchmarks
Guanhua Zhang, Moritz Hardt
We examine multi-task benchmarks in machine learning through the lens of social choice theory. We draw an analogy between benchmarks and electoral systems, where models are candida…
Stochastic Optimization for Performative Prediction
Celestine Mendler-Dünner, Juan C. Perdomo, Tijana Zrnic +1
In performative prediction, the choice of a model influences the distribution of future data, typically through actions taken based on the model's predictions. We initiate the stud…
Identity Crisis: Memorization and Generalization under Extreme Overparameterization
Chiyuan Zhang, Samy Bengio, Moritz Hardt +2
We study the interplay between memorization and generalization of overparameterized networks in the extreme case of a single training example and an identity-mapping task. We exami…
Computational Limits for Matrix Completion
Moritz Hardt, Raghu Meka, Prasad Raghavendra +1
Matrix Completion is the problem of recovering an unknown real-valued low-rank matrix from a subsample of its entries. Important recent results show that the problem can be solved…
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Mina Remeli, Moritz Hardt
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues…
How Benchmark Prediction from Fewer Data Misses the Mark
Guanhua Zhang, Florian E. Dorner, Moritz Hardt
Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets. Benchmark prediction (also cal…
Equality of Opportunity in Supervised Learning
Moritz Hardt, Eric Price, Nathan Srebro
We propose a criterion for discrimination against a specified sensitive attribute in supervised learning, where the goal is to predict some target based on available features. Assu…
A simple and practical algorithm for differentially private data release
Moritz Hardt, Katrina Ligett, Frank McSherry
We present new theoretical results on differentially private data release useful with respect to any target class of counting queries, coupled with experimental results on a variet…
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
Florian E. Dorner, Moritz Hardt
We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a giv…
Linear Dynamics: Clustering without identification
Chloe Ching-Yun Hsu, Michaela Hardt, Moritz Hardt
Linear dynamical systems are a fundamental and powerful parametric model class. However, identifying the parameters of a linear dynamical system is a venerable task, permitting pro…
Limits to Predicting Online Speech Using Large Language Models
Mina Remeli, Moritz Hardt, Robert C. Williamson
Our paper studies the predictability of online speech -- that is, how well language models learn to model the distribution of user generated content on X (previously Twitter). We d…
Private Data Release via Learning Thresholds
Moritz Hardt, Guy N. Rothblum, Rocco A. Servedio
This work considers computationally efficient privacy-preserving data release. We study the task of analyzing a database containing sensitive information about individual participa…
On the Geometry of Differential Privacy
Moritz Hardt, Kunal Talwar
We consider the noise complexity of differentially private mechanisms in the setting where the user asks linear queries $f\colon\Rn\to\Re$ non-adaptively. Here, the database is…
Scaling Open-Ended Reasoning to Predict the Future
Nikhil Chandak, Shashwat Goel, Ameya Prabhu +2
High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. T…
Evaluating language models as risk scores
André F. Cruz, Moritz Hardt, Celestine Mendler-Dünner
Current question-answering benchmarks predominantly focus on accuracy in realizable prediction tasks. Conditioned on a question and answer-key, does the most likely token match the…
Performative Power
Moritz Hardt, Meena Jagadeesan, Celestine Mendler-Dünner
We introduce the notion of performative power, which measures the ability of a firm operating an algorithmic system, such as a digital content recommendation platform, to cause cha…
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
Florian E. Dorner, Vivian Y. Nastl, Moritz Hardt
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefor…
Privately Releasing Conjunctions and the Statistical Query Barrier
Anupam Gupta, Moritz Hardt, Aaron Roth +1
Suppose we would like to know all answers to a set of statistical queries C on a data set up to small error, but we can only access the data itself using statistical queries. A tri…
Explaining an increase in predicted risk for clinical alerts
Michaela Hardt, Alvin Rajkomar, Gerardo Flores +5
Much work aims to explain a model's prediction on a static input. We consider explanations in a temporal setting where a stateful dynamical model produces a sequence of risk estima…
FutureSim: Replaying World Events to Evaluate Adaptive Agents
Shashwat Goel, Nikhil Chandak, Arvindh Arun +5
AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for rea…
Fast matrix completion without the condition number
Moritz Hardt, Mary Wootters
We give the first algorithm for Matrix Completion whose running time and sample complexity is polynomial in the rank of the unknown target matrix, linear in the dimension of the ma…
County-level Algorithmic Audit of Racial Bias in Twitter's Home Timeline
Luca Belli, Kyra Yee, Uthaipon Tantipongpipat +3
We report on the outcome of an audit of Twitter's Home Timeline ranking system. The goal of the audit was to determine if authors from some racial groups experience systematically…
What Makes ImageNet Look Unlike LAION
Ali Shirali, Moritz Hardt
ImageNet was famously created from Flickr image search results. What if we recreated ImageNet instead by searching the massive LAION dataset based on image captions alone? In this…
Retiring Adult: New Datasets for Fair Machine Learning
Frances Ding, Moritz Hardt, John Miller +1
Although the fairness community has recognized the importance of data, researchers in the area primarily rely on UCI Adult when it comes to tabular data. Derived from a 1994 US Cen…
Is your model predicting the past?
Moritz Hardt, Michael P. Kim
When does a machine learning model predict the future of individuals and when does it recite patterns that predate the individuals? In this work, we propose a distinction between t…
Balancing Competing Objectives with Noisy Data: Score-Based Classifiers for Welfare-Aware Machine Learning
Esther Rolf, Max Simchowitz, Sarah Dean +4
While real-world decisions involve many competing objectives, algorithmic decisions are often evaluated with a single objective function. In this paper, we study algorithmic polici…
Fairness Through Awareness
Cynthia Dwork, Moritz Hardt, Toniann Pitassi +2
We study fairness in classification, where individuals are classified, e.g., admitted to a university, and the goal is to prevent discrimination against individuals based on their…
Causal Inference out of Control: Estimating the Steerability of Consumption
Gary Cheng, Moritz Hardt, Celestine Mendler-Dünner
Regulators and academics are increasingly interested in the causal effect that algorithmic actions of a digital platform have on consumption. We introduce a general causal inferenc…
Identity Matters in Deep Learning
Moritz Hardt, Tengyu Ma
An emerging design principle in deep learning is that each layer of a deep artificial neural network should be able to easily express the identity transformation. This idea not onl…
Model Similarity Mitigates Test Set Overuse
Horia Mania, John Miller, Ludwig Schmidt +2
Excessive reuse of test data has become commonplace in today's machine learning workflows. Popular benchmarks, competitions, industrial scale tuning, among other applications, all…