papers

Publications (90)

stat.ME2017

Empirical Bayesian analysis of simultaneous changepoints in multiple data sequences

Zhou Fan, Lester Mackey

Copy number variations in cancer cells and volatility fluctuations in stock prices are commonly manifested as changepoints occurring at the same positions across related data seque…

stat.ML2020

Accurate Inference for Adaptive Linear Models

Yash Deshpande, Lester Mackey, Vasilis Syrgkanis +1

Estimators computed from adaptively collected data do not behave like their non-adaptive brethren. Rather, the sequential dependence of the collection policy can lead to severe dis…

stat.ML2023

Learning Rate Free Sampling in Constrained Domains

Louis Sharrock, Lester Mackey, Christopher Nemeth

We introduce a suite of new particle-based algorithms for sampling in constrained domains which are entirely learning rate free. Our approach leverages coin betting ideas from conv…

cs.SI2012

Jointly Predicting Links and Inferring Attributes using a Social-Attribute Network (SAN)

Neil Zhenqiang Gong, Ameet Talwalkar, Lester Mackey +6

The effects of social influence and homophily suggest that both network structure and node attribute information should inform the tasks of link prediction and node attribute infer…

cs.AI2026

Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner

Cai Zhou, Chenxiao Yang, Yi Hu +7

Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical and primary empirical results showing…

stat.ML2020

Measuring Sample Quality with Kernels

Jackson Gorham, Lester Mackey

Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these bia…

cs.LG2020

Teacher-Student Compression with Generative Adversarial Networks

Ruishan Liu, Nicolo Fusi, Lester Mackey

More accurate machine learning models often demand more computation and memory at test time, making them difficult to deploy on CPU- or memory-constrained devices. Teacher-student…

stat.ME2013

Combinatorial clustering and the beta negative binomial process

Tamara Broderick, Lester Mackey, John Paisley +1

We develop a Bayesian nonparametric approach to a general family of latent class problems in which individuals can belong simultaneously to multiple classes and where each class ca…

stat.ML2017

Improving Gibbs Sampler Scan Quality with DoGS

Ioannis Mitliagkas, Lester Mackey

The pairwise influence matrix of Dobrushin has long been used as an analytical tool to bound the rate of convergence of Gibbs sampling. In this work, we use Dobrushin influence as…

stat.CO2022

Scalable Spike-and-Slab

Niloy Biswas, Lester Mackey, Xiao-Li Meng

Spike-and-slab priors are commonly used for Bayesian variable selection, due to their interpretability and favorable statistical properties. However, existing samplers for spike-an…

cs.CL2024

Do Language Models Know When They're Hallucinating References?

Ayush Agrawal, Mirac Suzgun, Lester Mackey +1

State-of-the-art language models (LMs) are notoriously susceptible to generating hallucinated information. Such inaccurate outputs not only undermine the reliability of these model…

stat.ML2022

Initialization and Regularization of Factorized Neural Layers

Mikhail Khodak, Neil Tenenholtz, Lester Mackey +1

Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of…

stat.ML2025

Compress Then Test: Powerful Kernel Testing in Near-linear Time

Carles Domingo-Enrich, Raaz Dwivedi, Lester Mackey

Kernel two-sample testing provides a powerful framework for distinguishing any pair of distributions based on sample points. However, existing kernel tests either run in

cs.LG2026

WildCat: Near-Linear Attention in Theory and Practice

Tobias Schröder, Lester Mackey

We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks. While attention is a staple of modern network architectures, it…

stat.CO2018

Stein Points

Wilson Ye Chen, Lester Mackey, Jackson Gorham +2

An important task in computational statistics and machine learning is to approximate a posterior distribution with an empirical measure supported on a set of representative…

stat.AP2019

Improving Subseasonal Forecasting in the Western U.S. with Machine Learning

Jessica Hwang, Paulo Orenstein, Judah Cohen +2

Water managers in the western United States (U.S.) rely on longterm forecasts of temperature and precipitation to prepare for droughts and other wet weather extremes. To improve th…

stat.ML2024

Debiased Distribution Compression

Lingxiao Li, Raaz Dwivedi, Lester Mackey

Modern compression methods can summarize a target distribution more succinctly than i.i.d. sampling but require access to a low-bias input sequence like a Markov chain…

cs.CV2024

SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery

Konstantin Klemmer, Esther Rolf, Caleb Robinson +2

Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be…

stat.ML2024

Gradient Estimation with Discrete Stein Operators

Jiaxin Shi, Yuhao Zhou, Jessica Hwang +2

Gradient estimation -- approximating the gradient of an expectation with respect to the parameters of a distribution -- is central to the solution of many machine learning problems…

stat.ML2020

Cross-validation Confidence Intervals for Test Error

Pierre Bayle, Alexandre Bayle, Lucas Janson +1

This work develops central limit theorems for cross-validation and consistent estimators of its asymptotic variance under weak stability conditions on the learning algorithm. Toget…

econ.EM2020

Minimax Estimation of Conditional Moment Models

Nishanth Dikkala, Greg Lewis, Lester Mackey +1

We develop an approach for estimating models described via conditional moment restrictions, with a prototypical application being non-parametric instrumental variable regression. W…

stat.ML2026

The Relative Instability of Model Comparison with Cross-validation

Alexandre Bayle, Lucas Janson, Lester Mackey

Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surpris…

cs.LG2024

SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation

Mikhail Khodak, Lester Mackey, Alexandra Chouldechova +1

Disaggregated evaluation -- estimation of performance of a machine learning model on different subpopulations -- is a core task when assessing performance and group-fairness of AI…

stat.CO2020

Stein Point Markov Chain Monte Carlo

Wilson Ye Chen, Alessandro Barp, François-Xavier Briol +4

An important task in machine learning and statistics is the approximation of a probability measure by an empirical measure supported on a discrete point set. Stein Points are a cla…

math.OC2020

Importance Sampling via Local Sensitivity

Anant Raj, Cameron Musco, Lester Mackey

Given a loss function that can be written as the sum of losses over a large set of inputs , it is often desirable to approximate $…

stat.ML2015

Weighted Classification Cascades for Optimizing Discovery Significance in the HiggsML Challenge

Lester Mackey, Jordan Bryan, Man Yue Mo

We introduce a minorization-maximization approach to optimizing common measures of discovery significance in high energy physics. The approach alternates between solving a weighted…

cs.LG2026

Express Language Modeling

Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1

We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the sta…

cs.LG2026

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

Hannah Guan, Soukayna Mouatadid, Paulo Orenstein +10

Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes. Today, such forecasts enjoy unprecedented a…

math.OC2020

Accelerating Rescaled Gradient Descent: Fast Optimization of Smooth Functions

Ashia Wilson, Lester Mackey, Andre Wibisono

We present a family of algorithms, called descent algorithms, for optimizing convex and non-convex functions. We also introduce a new first-order algorithm, called rescaled gradien…

cs.AI2023

Reflections from the Workshop on AI-Assisted Decision Making for Conservation

Lily Xu, Esther Rolf, Sara Beery +21

In this white paper, we synthesize key points made during presentations and discussions from the AI-Assisted Decision Making for Conservation workshop, hosted by the Center for Res…

stat.ML2026

Low-Rank Thinning

Annabelle Michael Carrell, Albert Gong, Abhishek Shetty +2

The goal in thinning is to summarize a dataset using a small set of representative points. Remarkably, sub-Gaussian thinning algorithms like Kernel Halving and Compress can match t…

stat.ML2020

Stochastic Runge-Kutta Accelerates Langevin Monte Carlo and Beyond

Xuechen Li, Denny Wu, Lester Mackey +1

Sampling with Markov chain Monte Carlo methods often amounts to discretizing some continuous-time dynamics with numerical integration. In this paper, we establish the convergence r…

hep-ph2017

Jet-Images -- Deep Learning Edition

Luke de Oliveira, Michael Kagan, Lester Mackey +2

Building on the notion of a particle physics detector as a camera and the collimated streams of high energy particles, or jets, it measures as an image, we investigate the potentia…

stat.ML2026

Estimating Treatment Effects with Independent Component Analysis

Patrik Reizinger, Lester Mackey, Wieland Brendel +1

Independent Component Analysis (ICA) uses a measure of non-Gaussianity to identify latent sources from data and estimate their mixing coefficients (Shimizu et al., 2006). Meanwhile…

cs.LG2025

Informed Correctors for Discrete Diffusion Models

Yixiu Zhao, Jiaxin Shi, Feng Chen +3

Discrete diffusion has emerged as a powerful framework for generative modeling in discrete domains, yet efficiently sampling from these models remains challenging. Existing samplin…

math.PR2016

Multivariate Stein Factors for a Class of Strongly Log-concave Distributions

Lester Mackey, Jackson Gorham

We establish uniform bounds on the low-order derivatives of Stein equation solutions for a broad class of multivariate, strongly log-concave target distributions. These "Stein fact…

stat.CO2023

Bounding Wasserstein distance with couplings

Niloy Biswas, Lester Mackey

Markov chain Monte Carlo (MCMC) provides asymptotically consistent estimates of intractable posterior expectations as the number of iterations tends to infinity. However, in large…

cs.LG2021

Online Learning with Optimism and Delay

Genevieve Flaspohler, Francesco Orabona, Judah Cohen +4

Inspired by the demands of real-time climate and weather forecasting, we develop optimistic online learning algorithms that require no parameter tuning and have optimal regret guar…

stat.ME2022

Optimal Thinning of MCMC Output

Marina Riabiz, Wilson Chen, Jon Cockayne +4

The use of heuristics to assess the convergence and compress the output of Markov chain Monte Carlo can be sub-optimal in terms of the empirical approximations that are produced. T…

stat.ML2020

Stochastic Stein Discrepancies

Jackson Gorham, Anant Raj, Lester Mackey

Stein discrepancies (SDs) monitor convergence and non-convergence in approximate inference when exact integration and sampling are intractable. However, the computation of a Stein…

stat.ML2020

Approximate Cross-validation: Guarantees for Model Assessment and Selection

Ashia Wilson, Maximilian Kasy, Lester Mackey

Cross-validation (CV) is a popular approach for assessing and selecting predictive models. However, when the number of folds is large, CV suffers from a need to repeatedly refit a…

math.PR2013

Deriving Matrix Concentration Inequalities from Kernel Couplings

Daniel Paulin, Lester Mackey, Joel A. Tropp

This paper derives exponential tail bounds and polynomial moment inequalities for the spectral norm deviation of a random matrix from its mean value. The argument depends on a matr…

cs.LG2025

KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning

Eric Zimmermann, Harley Wiltzer, Justin Szeto +2

Recent breakthroughs in self-supervised Joint-Embedding Predictive Architectures (JEPAs) have established that regularizing Euclidean representations toward isotropic Gaussian prio…

stat.ME2023

Should I Stop or Should I Go: Early Stopping with Heterogeneous Populations

Hammaad Adam, Fan Yin, Huibin +5

Randomized experiments often need to be stopped prematurely due to the treatment having an unintended harmful effect. Existing methods that determine when to stop an experiment ear…

cs.LG2026

Integrating chemical structures as treatments improves representations of microscopy images for morphological profiling

Yemin Yu, Emre Hayir, Neil Tenenholtz +5

Recent advances in self-supervised deep learning have improved our ability to quantify cellular morphological changes in high-throughput microscopy screens, a process known as morp…

cs.CV2013

Distributed Low-rank Subspace Segmentation

Ameet Talwalkar, Lester Mackey, Yadong Mu +2

Vision problems ranging from image clustering to motion segmentation to semi-supervised learning can naturally be framed as subspace segmentation problems, in which one aims to rec…

cs.LG2022

Social Norm Bias: Residual Harms of Fairness-Aware Algorithms

Myra Cheng, Maria De-Arteaga, Lester Mackey +1

Many modern machine learning algorithms mitigate bias by enforcing fairness constraints across coarsely-defined groups related to a sensitive attribute like gender or race. However…

cs.LG2026

Thinned Mean Field Langevin Dynamics

Zonghao Chen, Heishiro Kanagawa, François-Xavier Briol +2

Several important learning tasks can be formulated as minimizing an entropy-regularized objective over an appropriate space of probability distributions. Mean-field Langevin dynami…

cs.CV2018

Expert identification of visual primitives used by CNNs during mammogram classification

Jimmy Wu, Diondra Peck, Scott Hsieh +6

This work interprets the internal representations of deep neural networks trained for classification of diseased tissue in 2D mammograms. We propose an expert-in-the-loop interpret…

math.PR2026

Efficient Concentration with Gaussian Approximation

Morgane Austern, Lester Mackey

Concentration inequalities for the sample mean, like those due to Bernstein, Hoeffding, and Bentkus, are valid for any sample size but overly conservative, yielding confidence inte…

stat.ML2020

Single Point Transductive Prediction

Nilesh Tripuraneni, Lester Mackey

Standard methods in supervised learning separate training and prediction: the model is fit independently of any test points it may encounter. However, can knowledge of the next tes…

stat.ML2018

Measuring Sample Quality with Diffusions

Jackson Gorham, Andrew B. Duncan, Sebastian J. Vollmer +1

Stein's method for measuring convergence to a continuous target distribution relies on an operator characterizing the target and Stein factor bounds on the solutions of an associat…

stat.ML2022

Sampling with Mirrored Stein Operators

Jiaxin Shi, Chang Liu, Lester Mackey

We introduce a new family of particle evolution samplers suitable for constrained domains and non-Euclidean geometries. Stein Variational Mirror Descent and Mirrored Stein Variatio…

math.ST2023

Near-optimal inference in adaptive linear regression

Koulik Khamaru, Yash Deshpande, Tor Lattimore +2

When data is collected in an adaptive manner, even simple methods like ordinary least squares can exhibit non-normal asymptotic behavior. As an undesirable consequence, hypothesis…

stat.ME2023

Independent finite approximations for Bayesian nonparametric inference

Tin D. Nguyen, Jonathan Huggins, Lorenzo Masoero +2

Completely random measures (CRMs) and their normalizations (NCRMs) offer flexible models in Bayesian nonparametrics. But their infinite dimensionality presents challenges for infer…

math.ST2025

Cheap Permutation Testing

Carles Domingo-Enrich, Raaz Dwivedi, Lester Mackey

Permutation tests are a popular choice for distinguishing distributions and testing independence, due to their exact, finite-sample control of false positives and their minimax opt…

cs.LG2023

Adaptive Bias Correction for Improved Subseasonal Forecasting

Soukayna Mouatadid, Paulo Orenstein, Genevieve Flaspohler +4

Subseasonal forecasting -- predicting temperature and precipitation 2 to 6 weeks ahead -- is critical for effective water allocation, wildfire management, and drought and flood mit…

math.PR2014

Matrix concentration inequalities via the method of exchangeable pairs

Lester Mackey, Michael I. Jordan, Richard Y. Chen +2

This paper derives exponential concentration inequalities and polynomial moment inequalities for the spectral norm of a random matrix. The analysis requires a matrix extension of t…

cs.LG2022

Budget-Constrained Bounds for Mini-Batch Estimation of Optimal Transport

David Alvarez-Melis, Nicolò Fusi, Lester Mackey +1

Optimal Transport (OT) is a fundamental tool for comparing probability distributions, but its exact computation remains prohibitive for large datasets. In this work, we introduce n…

stat.ML2023

A Kernel Stein Test for Comparing Latent Variable Models

Heishiro Kanagawa, Wittawat Jitkrittum, Lester Mackey +2

We propose a kernel-based nonparametric test of relative goodness of fit, where the goal is to compare two models, both of which may have unobserved latent variables, such that the…

stat.ME2022

Stein's Method Meets Computational Statistics: A Review of Some Recent Developments

Andreas Anastasiou, Alessandro Barp, François-Xavier Briol +11

Stein's method compares probability distributions through the study of a class of linear operators called Stein operators. While mainly studied in probability and used to underpin…

math.PR2025

Bounding Hellinger Distance with Stein's Method

Morgane Austern, Lester Mackey

This work introduces a new, explicit bound on the Hellinger distance between a continuous random variable and a Gaussian with matching mean and variance. As example applications, w…

stat.ML2024

Kernel Thinning

Raaz Dwivedi, Lester Mackey

We introduce kernel thinning, a new procedure for compressing a distribution more effectively than i.i.d. sampling or standard thinning. Given a suitable reproducing k…

cs.LG2020

Model-specific Data Subsampling with Influence Functions

Anant Raj, Cameron Musco, Lester Mackey +1

Model selection requires repeatedly evaluating models on a given dataset and measuring their relative performances. In modern applications of machine learning, the models being con…

cs.LG2009

Feature-Weighted Linear Stacking

Joseph Sill, Gabor Takacs, Lester Mackey +1

Ensemble methods, such as stacking, are designed to boost predictive accuracy by blending the predictions of multiple machine learning models. Recent work has shown that the use of…

stat.ML2025

Generalized Kernel Thinning

Raaz Dwivedi, Lester Mackey

The kernel thinning (KT) algorithm of Dwivedi and Mackey (2021) compresses a probability distribution more effectively than independent sampling by targeting a reproducing kernel H…

stat.ML2020

Weighted Meta-Learning

Diana Cai, Rishit Sheth, Lester Mackey +1

Meta-learning leverages related source tasks to learn an initialization that can be quickly fine-tuned to a target task with limited labeled examples. However, many popular meta-le…

math.ST2022

Minimum Stein Discrepancy Estimators

Alessandro Barp, Francois-Xavier Briol, Andrew B. Duncan +2

When maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates. We…

stat.ML2025

It's Hard to Be Normal: The Impact of Noise on Structure-agnostic Estimation

Jikai Jin, Lester Mackey, Vasilis Syrgkanis

Structure-agnostic causal inference studies how well one can estimate a treatment effect given black-box machine learning estimates of nuisance functions (like the impact of confou…

stat.ML2025

Targeted Separation and Convergence with Kernel Discrepancies

Alessandro Barp, Carl-Johann Simon-Gabriel, Mark Girolami +1

Maximum mean discrepancies (MMDs) like the kernel Stein discrepancy (KSD) have grown central to a wide range of applications, including hypothesis testing, sampler selection, distr…

physics.ao-ph2024

SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and Benchmarking

Soukayna Mouatadid, Paulo Orenstein, Genevieve Flaspohler +8

Subseasonal forecasting of the weather two to six weeks in advance is critical for resource allocation and advance disaster notice but poses many challenges for the forecasting com…

cs.LG2013

Distributed Matrix Completion and Robust Factorization

Lester Mackey, Ameet Talwalkar, Michael I. Jordan

If learning methods are to scale to the massive sizes of modern datasets, it is essential for the field of machine learning to embrace parallel and distributed computing. Inspired…

hep-ph2015

Fuzzy Jets

Lester Mackey, Benjamin Nachman, Ariel Schwartzman +1

Collimated streams of particles produced in high energy physics experiments are organized using clustering algorithms to form jets. To construct jets, the experimental collaboratio…

cs.LG2023

A Finite-Particle Convergence Rate for Stein Variational Gradient Descent

Jiaxin Shi, Lester Mackey

We provide the first finite-particle convergence rate for Stein variational gradient descent (SVGD), a popular algorithm for approximating a probability distribution with a collect…

stat.ML2025

Controlling Moments with Kernel Stein Discrepancies

Heishiro Kanagawa, Alessandro Barp, Arthur Gretton +1

Kernel Stein discrepancies (KSDs) measure the quality of a distributional approximation and can be computed even when the target density has an intractable normalizing constant. No…

cs.CV2021

DeepMiner: Discovering Interpretable Representations for Mammogram Classification and Explanation

Jimmy Wu, Bolei Zhou, Diondra Peck +4

We propose DeepMiner, a framework to discover interpretable representations in deep neural networks and to build explanations for medical predictions. By probing convolutional neur…

cs.LG2021

Metrizing Weak Convergence with Maximum Mean Discrepancies

Carl-Johann Simon-Gabriel, Alessandro Barp, Bernhard Schölkopf +1

This paper characterizes the maximum mean discrepancies (MMD) that metrize the weak convergence of probability measures for a wide class of kernels. More precisely, we prove that,…

stat.ML2019

Measuring Sample Quality with Stein's Method

Jackson Gorham, Lester Mackey

To improve the efficiency of Monte Carlo estimation, practitioners are turning to biased Markov chain Monte Carlo procedures that trade off asymptotic exactness for computational s…

cs.LG2026

Domain-Aware Scaling Laws Uncover Data Synergy

Kimia Hamidieh, Lester Mackey, David Alvarez-Melis

The paper defines and measures how mixing data from different domains during language model pretraining can produce synergistic or interfering effects, and shows that accounting fo…

#data synergy#scaling laws#language model pretraining#domain mixtures
cs.CL2024

Adapting Language Models via Token Translation

Zhili Feng, Tanya Marwah, Nicolo Fusi +2

Modern large language models use a fixed tokenizer to effectively compress text drawn from a source domain. However, applying the same tokenizer to a new target domain often leads…

stat.ML2026

Probabilistic Inference and Learning with Stein's Method

Qiang Liu, Lester Mackey, Chris Oates

This monograph provides a rigorous overview of theoretical and methodological aspects of probabilistic inference and learning with Stein's method. Recipes are provided for construc…

stat.ML2022

Distribution Compression in Near-linear Time

Abhishek Shetty, Raaz Dwivedi, Lester Mackey

In distribution compression, one aims to accurately summarize a probability distribution using a small number of representative points. Near-optimal thinning procedure…

math.ST2013

The asymptotics of ranking algorithms

John C. Duchi, Lester Mackey, Michael I. Jordan

We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical proced…

cs.IT2014

Corrupted Sensing: Novel Guarantees for Separating Structured Signals

Rina Foygel, Lester Mackey

We study the problem of corrupted sensing, a generalization of compressed sensing in which one aims to recover a signal from a collection of corrupted or unreliable measurements. W…

cs.CL2024

Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation

Eric Zelikman, Eliana Lorch, Lester Mackey +1

Several recent advances in AI systems solve problems by providing a "scaffolding" program that structures multiple calls to language models (LMs) to generate better outputs. A scaf…

cs.LG2018

Orthogonal Machine Learning: Power and Limitations

Lester Mackey, Vasilis Syrgkanis, Ilias Zadik

Double machine learning provides -consistent estimates of parameters of interest even when high-dimensional or nonparametric nuisance parameters are estimated at an $n^{-…

stat.ML2021

Random Feature Stein Discrepancies

Jonathan H. Huggins, Lester Mackey

Computable Stein discrepancies have been deployed for a variety of applications, ranging from sampler selection in posterior inference to approximate Bayesian inference to goodness…

stat.ML2019

Global Non-convex Optimization with Discretized Diffusions

Murat A. Erdogdu, Lester Mackey, Ohad Shamir

An Euler discretization of the Langevin diffusion is known to converge to the global minimizers of certain convex and non-convex optimization problems. We show that this property h…

math.PR2014

Efron-Stein Inequalities for Random Matrices

Daniel Paulin, Lester Mackey, Joel A. Tropp

This paper establishes new concentration inequalities for random matrices constructed from independent random variables. These results are analogous with the generalized Efron-Stei…

stat.ML2021

Knowledge Distillation as Semiparametric Inference

Tri Dao, Govinda M Kamath, Vasilis Syrgkanis +1

A popular approach to model compression is to train an inexpensive student model to mimic the class probabilities of a highly accurate but cumbersome teacher model. Surprisingly, t…