Publications (90)
Empirical Bayesian analysis of simultaneous changepoints in multiple data sequences
Zhou Fan, Lester Mackey
Copy number variations in cancer cells and volatility fluctuations in stock prices are commonly manifested as changepoints occurring at the same positions across related data seque…
Accurate Inference for Adaptive Linear Models
Yash Deshpande, Lester Mackey, Vasilis Syrgkanis +1
Estimators computed from adaptively collected data do not behave like their non-adaptive brethren. Rather, the sequential dependence of the collection policy can lead to severe dis…
Learning Rate Free Sampling in Constrained Domains
Louis Sharrock, Lester Mackey, Christopher Nemeth
We introduce a suite of new particle-based algorithms for sampling in constrained domains which are entirely learning rate free. Our approach leverages coin betting ideas from conv…
Jointly Predicting Links and Inferring Attributes using a Social-Attribute Network (SAN)
Neil Zhenqiang Gong, Ameet Talwalkar, Lester Mackey +6
The effects of social influence and homophily suggest that both network structure and node attribute information should inform the tasks of link prediction and node attribute infer…
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
Cai Zhou, Chenxiao Yang, Yi Hu +7
Diffusion language models, especially masked discrete diffusion models, have achieved great success recently. While there are some theoretical and primary empirical results showing…
Measuring Sample Quality with Kernels
Jackson Gorham, Lester Mackey
Approximate Markov chain Monte Carlo (MCMC) offers the promise of more rapid sampling at the cost of more biased inference. Since standard MCMC diagnostics fail to detect these bia…
Teacher-Student Compression with Generative Adversarial Networks
Ruishan Liu, Nicolo Fusi, Lester Mackey
More accurate machine learning models often demand more computation and memory at test time, making them difficult to deploy on CPU- or memory-constrained devices. Teacher-student…
Combinatorial clustering and the beta negative binomial process
Tamara Broderick, Lester Mackey, John Paisley +1
We develop a Bayesian nonparametric approach to a general family of latent class problems in which individuals can belong simultaneously to multiple classes and where each class ca…
Improving Gibbs Sampler Scan Quality with DoGS
Ioannis Mitliagkas, Lester Mackey
The pairwise influence matrix of Dobrushin has long been used as an analytical tool to bound the rate of convergence of Gibbs sampling. In this work, we use Dobrushin influence as…
Scalable Spike-and-Slab
Niloy Biswas, Lester Mackey, Xiao-Li Meng
Spike-and-slab priors are commonly used for Bayesian variable selection, due to their interpretability and favorable statistical properties. However, existing samplers for spike-an…
Do Language Models Know When They're Hallucinating References?
Ayush Agrawal, Mirac Suzgun, Lester Mackey +1
State-of-the-art language models (LMs) are notoriously susceptible to generating hallucinated information. Such inaccurate outputs not only undermine the reliability of these model…
Initialization and Regularization of Factorized Neural Layers
Mikhail Khodak, Neil Tenenholtz, Lester Mackey +1
Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of…
Compress Then Test: Powerful Kernel Testing in Near-linear Time
Carles Domingo-Enrich, Raaz Dwivedi, Lester Mackey
Kernel two-sample testing provides a powerful framework for distinguishing any pair of distributions based on sample points. However, existing kernel tests either run in …
WildCat: Near-Linear Attention in Theory and Practice
Tobias Schröder, Lester Mackey
We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks. While attention is a staple of modern network architectures, it…
Stein Points
Wilson Ye Chen, Lester Mackey, Jackson Gorham +2
An important task in computational statistics and machine learning is to approximate a posterior distribution with an empirical measure supported on a set of representative…
Improving Subseasonal Forecasting in the Western U.S. with Machine Learning
Jessica Hwang, Paulo Orenstein, Judah Cohen +2
Water managers in the western United States (U.S.) rely on longterm forecasts of temperature and precipitation to prepare for droughts and other wet weather extremes. To improve th…
Debiased Distribution Compression
Lingxiao Li, Raaz Dwivedi, Lester Mackey
Modern compression methods can summarize a target distribution more succinctly than i.i.d. sampling but require access to a low-bias input sequence like a Markov chain…
SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
Konstantin Klemmer, Esther Rolf, Caleb Robinson +2
Geographic information is essential for modeling tasks in fields ranging from ecology to epidemiology. However, extracting relevant location characteristics for a given task can be…
Gradient Estimation with Discrete Stein Operators
Jiaxin Shi, Yuhao Zhou, Jessica Hwang +2
Gradient estimation -- approximating the gradient of an expectation with respect to the parameters of a distribution -- is central to the solution of many machine learning problems…
Cross-validation Confidence Intervals for Test Error
Pierre Bayle, Alexandre Bayle, Lucas Janson +1
This work develops central limit theorems for cross-validation and consistent estimators of its asymptotic variance under weak stability conditions on the learning algorithm. Toget…
Minimax Estimation of Conditional Moment Models
Nishanth Dikkala, Greg Lewis, Lester Mackey +1
We develop an approach for estimating models described via conditional moment restrictions, with a prototypical application being non-parametric instrumental variable regression. W…
The Relative Instability of Model Comparison with Cross-validation
Alexandre Bayle, Lucas Janson, Lester Mackey
Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surpris…
SureMap: Simultaneous Mean Estimation for Single-Task and Multi-Task Disaggregated Evaluation
Mikhail Khodak, Lester Mackey, Alexandra Chouldechova +1
Disaggregated evaluation -- estimation of performance of a machine learning model on different subpopulations -- is a core task when assessing performance and group-fairness of AI…
Stein Point Markov Chain Monte Carlo
Wilson Ye Chen, Alessandro Barp, François-Xavier Briol +4
An important task in machine learning and statistics is the approximation of a probability measure by an empirical measure supported on a discrete point set. Stein Points are a cla…
Importance Sampling via Local Sensitivity
Anant Raj, Cameron Musco, Lester Mackey
Given a loss function that can be written as the sum of losses over a large set of inputs , it is often desirable to approximate $…
Weighted Classification Cascades for Optimizing Discovery Significance in the HiggsML Challenge
Lester Mackey, Jordan Bryan, Man Yue Mo
We introduce a minorization-maximization approach to optimizing common measures of discovery significance in high energy physics. The approach alternates between solving a weighted…
Express Language Modeling
Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi +1
We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the sta…
Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction
Hannah Guan, Soukayna Mouatadid, Paulo Orenstein +10
Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes. Today, such forecasts enjoy unprecedented a…
Accelerating Rescaled Gradient Descent: Fast Optimization of Smooth Functions
Ashia Wilson, Lester Mackey, Andre Wibisono
We present a family of algorithms, called descent algorithms, for optimizing convex and non-convex functions. We also introduce a new first-order algorithm, called rescaled gradien…
Reflections from the Workshop on AI-Assisted Decision Making for Conservation
Lily Xu, Esther Rolf, Sara Beery +21
In this white paper, we synthesize key points made during presentations and discussions from the AI-Assisted Decision Making for Conservation workshop, hosted by the Center for Res…
Low-Rank Thinning
Annabelle Michael Carrell, Albert Gong, Abhishek Shetty +2
The goal in thinning is to summarize a dataset using a small set of representative points. Remarkably, sub-Gaussian thinning algorithms like Kernel Halving and Compress can match t…
Stochastic Runge-Kutta Accelerates Langevin Monte Carlo and Beyond
Xuechen Li, Denny Wu, Lester Mackey +1
Sampling with Markov chain Monte Carlo methods often amounts to discretizing some continuous-time dynamics with numerical integration. In this paper, we establish the convergence r…
Jet-Images -- Deep Learning Edition
Luke de Oliveira, Michael Kagan, Lester Mackey +2
Building on the notion of a particle physics detector as a camera and the collimated streams of high energy particles, or jets, it measures as an image, we investigate the potentia…
Estimating Treatment Effects with Independent Component Analysis
Patrik Reizinger, Lester Mackey, Wieland Brendel +1
Independent Component Analysis (ICA) uses a measure of non-Gaussianity to identify latent sources from data and estimate their mixing coefficients (Shimizu et al., 2006). Meanwhile…
Informed Correctors for Discrete Diffusion Models
Yixiu Zhao, Jiaxin Shi, Feng Chen +3
Discrete diffusion has emerged as a powerful framework for generative modeling in discrete domains, yet efficiently sampling from these models remains challenging. Existing samplin…
Multivariate Stein Factors for a Class of Strongly Log-concave Distributions
Lester Mackey, Jackson Gorham
We establish uniform bounds on the low-order derivatives of Stein equation solutions for a broad class of multivariate, strongly log-concave target distributions. These "Stein fact…
Bounding Wasserstein distance with couplings
Niloy Biswas, Lester Mackey
Markov chain Monte Carlo (MCMC) provides asymptotically consistent estimates of intractable posterior expectations as the number of iterations tends to infinity. However, in large…
Online Learning with Optimism and Delay
Genevieve Flaspohler, Francesco Orabona, Judah Cohen +4
Inspired by the demands of real-time climate and weather forecasting, we develop optimistic online learning algorithms that require no parameter tuning and have optimal regret guar…
Optimal Thinning of MCMC Output
Marina Riabiz, Wilson Chen, Jon Cockayne +4
The use of heuristics to assess the convergence and compress the output of Markov chain Monte Carlo can be sub-optimal in terms of the empirical approximations that are produced. T…
Stochastic Stein Discrepancies
Jackson Gorham, Anant Raj, Lester Mackey
Stein discrepancies (SDs) monitor convergence and non-convergence in approximate inference when exact integration and sampling are intractable. However, the computation of a Stein…
Approximate Cross-validation: Guarantees for Model Assessment and Selection
Ashia Wilson, Maximilian Kasy, Lester Mackey
Cross-validation (CV) is a popular approach for assessing and selecting predictive models. However, when the number of folds is large, CV suffers from a need to repeatedly refit a…
Deriving Matrix Concentration Inequalities from Kernel Couplings
Daniel Paulin, Lester Mackey, Joel A. Tropp
This paper derives exponential tail bounds and polynomial moment inequalities for the spectral norm deviation of a random matrix from its mean value. The argument depends on a matr…
KerJEPA: Kernel Discrepancies for Euclidean Self-Supervised Learning
Eric Zimmermann, Harley Wiltzer, Justin Szeto +2
Recent breakthroughs in self-supervised Joint-Embedding Predictive Architectures (JEPAs) have established that regularizing Euclidean representations toward isotropic Gaussian prio…
Should I Stop or Should I Go: Early Stopping with Heterogeneous Populations
Hammaad Adam, Fan Yin, Huibin +5
Randomized experiments often need to be stopped prematurely due to the treatment having an unintended harmful effect. Existing methods that determine when to stop an experiment ear…
Integrating chemical structures as treatments improves representations of microscopy images for morphological profiling
Yemin Yu, Emre Hayir, Neil Tenenholtz +5
Recent advances in self-supervised deep learning have improved our ability to quantify cellular morphological changes in high-throughput microscopy screens, a process known as morp…
Distributed Low-rank Subspace Segmentation
Ameet Talwalkar, Lester Mackey, Yadong Mu +2
Vision problems ranging from image clustering to motion segmentation to semi-supervised learning can naturally be framed as subspace segmentation problems, in which one aims to rec…
Social Norm Bias: Residual Harms of Fairness-Aware Algorithms
Myra Cheng, Maria De-Arteaga, Lester Mackey +1
Many modern machine learning algorithms mitigate bias by enforcing fairness constraints across coarsely-defined groups related to a sensitive attribute like gender or race. However…
Thinned Mean Field Langevin Dynamics
Zonghao Chen, Heishiro Kanagawa, François-Xavier Briol +2
Several important learning tasks can be formulated as minimizing an entropy-regularized objective over an appropriate space of probability distributions. Mean-field Langevin dynami…
Expert identification of visual primitives used by CNNs during mammogram classification
Jimmy Wu, Diondra Peck, Scott Hsieh +6
This work interprets the internal representations of deep neural networks trained for classification of diseased tissue in 2D mammograms. We propose an expert-in-the-loop interpret…
Efficient Concentration with Gaussian Approximation
Morgane Austern, Lester Mackey
Concentration inequalities for the sample mean, like those due to Bernstein, Hoeffding, and Bentkus, are valid for any sample size but overly conservative, yielding confidence inte…
Single Point Transductive Prediction
Nilesh Tripuraneni, Lester Mackey
Standard methods in supervised learning separate training and prediction: the model is fit independently of any test points it may encounter. However, can knowledge of the next tes…
Measuring Sample Quality with Diffusions
Jackson Gorham, Andrew B. Duncan, Sebastian J. Vollmer +1
Stein's method for measuring convergence to a continuous target distribution relies on an operator characterizing the target and Stein factor bounds on the solutions of an associat…
Sampling with Mirrored Stein Operators
Jiaxin Shi, Chang Liu, Lester Mackey
We introduce a new family of particle evolution samplers suitable for constrained domains and non-Euclidean geometries. Stein Variational Mirror Descent and Mirrored Stein Variatio…
Near-optimal inference in adaptive linear regression
Koulik Khamaru, Yash Deshpande, Tor Lattimore +2
When data is collected in an adaptive manner, even simple methods like ordinary least squares can exhibit non-normal asymptotic behavior. As an undesirable consequence, hypothesis…
Independent finite approximations for Bayesian nonparametric inference
Tin D. Nguyen, Jonathan Huggins, Lorenzo Masoero +2
Completely random measures (CRMs) and their normalizations (NCRMs) offer flexible models in Bayesian nonparametrics. But their infinite dimensionality presents challenges for infer…
Cheap Permutation Testing
Carles Domingo-Enrich, Raaz Dwivedi, Lester Mackey
Permutation tests are a popular choice for distinguishing distributions and testing independence, due to their exact, finite-sample control of false positives and their minimax opt…
Adaptive Bias Correction for Improved Subseasonal Forecasting
Soukayna Mouatadid, Paulo Orenstein, Genevieve Flaspohler +4
Subseasonal forecasting -- predicting temperature and precipitation 2 to 6 weeks ahead -- is critical for effective water allocation, wildfire management, and drought and flood mit…
Matrix concentration inequalities via the method of exchangeable pairs
Lester Mackey, Michael I. Jordan, Richard Y. Chen +2
This paper derives exponential concentration inequalities and polynomial moment inequalities for the spectral norm of a random matrix. The analysis requires a matrix extension of t…
Budget-Constrained Bounds for Mini-Batch Estimation of Optimal Transport
David Alvarez-Melis, Nicolò Fusi, Lester Mackey +1
Optimal Transport (OT) is a fundamental tool for comparing probability distributions, but its exact computation remains prohibitive for large datasets. In this work, we introduce n…
A Kernel Stein Test for Comparing Latent Variable Models
Heishiro Kanagawa, Wittawat Jitkrittum, Lester Mackey +2
We propose a kernel-based nonparametric test of relative goodness of fit, where the goal is to compare two models, both of which may have unobserved latent variables, such that the…
Stein's Method Meets Computational Statistics: A Review of Some Recent Developments
Andreas Anastasiou, Alessandro Barp, François-Xavier Briol +11
Stein's method compares probability distributions through the study of a class of linear operators called Stein operators. While mainly studied in probability and used to underpin…
Bounding Hellinger Distance with Stein's Method
Morgane Austern, Lester Mackey
This work introduces a new, explicit bound on the Hellinger distance between a continuous random variable and a Gaussian with matching mean and variance. As example applications, w…
Kernel Thinning
Raaz Dwivedi, Lester Mackey
We introduce kernel thinning, a new procedure for compressing a distribution more effectively than i.i.d. sampling or standard thinning. Given a suitable reproducing k…
Model-specific Data Subsampling with Influence Functions
Anant Raj, Cameron Musco, Lester Mackey +1
Model selection requires repeatedly evaluating models on a given dataset and measuring their relative performances. In modern applications of machine learning, the models being con…
Feature-Weighted Linear Stacking
Joseph Sill, Gabor Takacs, Lester Mackey +1
Ensemble methods, such as stacking, are designed to boost predictive accuracy by blending the predictions of multiple machine learning models. Recent work has shown that the use of…
Generalized Kernel Thinning
Raaz Dwivedi, Lester Mackey
The kernel thinning (KT) algorithm of Dwivedi and Mackey (2021) compresses a probability distribution more effectively than independent sampling by targeting a reproducing kernel H…
Weighted Meta-Learning
Diana Cai, Rishit Sheth, Lester Mackey +1
Meta-learning leverages related source tasks to learn an initialization that can be quickly fine-tuned to a target task with limited labeled examples. However, many popular meta-le…
Minimum Stein Discrepancy Estimators
Alessandro Barp, Francois-Xavier Briol, Andrew B. Duncan +2
When maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates. We…
It's Hard to Be Normal: The Impact of Noise on Structure-agnostic Estimation
Jikai Jin, Lester Mackey, Vasilis Syrgkanis
Structure-agnostic causal inference studies how well one can estimate a treatment effect given black-box machine learning estimates of nuisance functions (like the impact of confou…
Targeted Separation and Convergence with Kernel Discrepancies
Alessandro Barp, Carl-Johann Simon-Gabriel, Mark Girolami +1
Maximum mean discrepancies (MMDs) like the kernel Stein discrepancy (KSD) have grown central to a wide range of applications, including hypothesis testing, sampler selection, distr…
SubseasonalClimateUSA: A Dataset for Subseasonal Forecasting and Benchmarking
Soukayna Mouatadid, Paulo Orenstein, Genevieve Flaspohler +8
Subseasonal forecasting of the weather two to six weeks in advance is critical for resource allocation and advance disaster notice but poses many challenges for the forecasting com…
Distributed Matrix Completion and Robust Factorization
Lester Mackey, Ameet Talwalkar, Michael I. Jordan
If learning methods are to scale to the massive sizes of modern datasets, it is essential for the field of machine learning to embrace parallel and distributed computing. Inspired…
Fuzzy Jets
Lester Mackey, Benjamin Nachman, Ariel Schwartzman +1
Collimated streams of particles produced in high energy physics experiments are organized using clustering algorithms to form jets. To construct jets, the experimental collaboratio…
A Finite-Particle Convergence Rate for Stein Variational Gradient Descent
Jiaxin Shi, Lester Mackey
We provide the first finite-particle convergence rate for Stein variational gradient descent (SVGD), a popular algorithm for approximating a probability distribution with a collect…
Controlling Moments with Kernel Stein Discrepancies
Heishiro Kanagawa, Alessandro Barp, Arthur Gretton +1
Kernel Stein discrepancies (KSDs) measure the quality of a distributional approximation and can be computed even when the target density has an intractable normalizing constant. No…
DeepMiner: Discovering Interpretable Representations for Mammogram Classification and Explanation
Jimmy Wu, Bolei Zhou, Diondra Peck +4
We propose DeepMiner, a framework to discover interpretable representations in deep neural networks and to build explanations for medical predictions. By probing convolutional neur…
Metrizing Weak Convergence with Maximum Mean Discrepancies
Carl-Johann Simon-Gabriel, Alessandro Barp, Bernhard Schölkopf +1
This paper characterizes the maximum mean discrepancies (MMD) that metrize the weak convergence of probability measures for a wide class of kernels. More precisely, we prove that,…
Measuring Sample Quality with Stein's Method
Jackson Gorham, Lester Mackey
To improve the efficiency of Monte Carlo estimation, practitioners are turning to biased Markov chain Monte Carlo procedures that trade off asymptotic exactness for computational s…
Domain-Aware Scaling Laws Uncover Data Synergy
Kimia Hamidieh, Lester Mackey, David Alvarez-Melis
The paper defines and measures how mixing data from different domains during language model pretraining can produce synergistic or interfering effects, and shows that accounting fo…
Adapting Language Models via Token Translation
Zhili Feng, Tanya Marwah, Nicolo Fusi +2
Modern large language models use a fixed tokenizer to effectively compress text drawn from a source domain. However, applying the same tokenizer to a new target domain often leads…
Probabilistic Inference and Learning with Stein's Method
Qiang Liu, Lester Mackey, Chris Oates
This monograph provides a rigorous overview of theoretical and methodological aspects of probabilistic inference and learning with Stein's method. Recipes are provided for construc…
Distribution Compression in Near-linear Time
Abhishek Shetty, Raaz Dwivedi, Lester Mackey
In distribution compression, one aims to accurately summarize a probability distribution using a small number of representative points. Near-optimal thinning procedure…
The asymptotics of ranking algorithms
John C. Duchi, Lester Mackey, Michael I. Jordan
We consider the predictive problem of supervised ranking, where the task is to rank sets of candidate items returned in response to queries. Although there exist statistical proced…
Corrupted Sensing: Novel Guarantees for Separating Structured Signals
Rina Foygel, Lester Mackey
We study the problem of corrupted sensing, a generalization of compressed sensing in which one aims to recover a signal from a collection of corrupted or unreliable measurements. W…
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
Eric Zelikman, Eliana Lorch, Lester Mackey +1
Several recent advances in AI systems solve problems by providing a "scaffolding" program that structures multiple calls to language models (LMs) to generate better outputs. A scaf…
Orthogonal Machine Learning: Power and Limitations
Lester Mackey, Vasilis Syrgkanis, Ilias Zadik
Double machine learning provides -consistent estimates of parameters of interest even when high-dimensional or nonparametric nuisance parameters are estimated at an $n^{-…
Random Feature Stein Discrepancies
Jonathan H. Huggins, Lester Mackey
Computable Stein discrepancies have been deployed for a variety of applications, ranging from sampler selection in posterior inference to approximate Bayesian inference to goodness…
Global Non-convex Optimization with Discretized Diffusions
Murat A. Erdogdu, Lester Mackey, Ohad Shamir
An Euler discretization of the Langevin diffusion is known to converge to the global minimizers of certain convex and non-convex optimization problems. We show that this property h…
Efron-Stein Inequalities for Random Matrices
Daniel Paulin, Lester Mackey, Joel A. Tropp
This paper establishes new concentration inequalities for random matrices constructed from independent random variables. These results are analogous with the generalized Efron-Stei…
Knowledge Distillation as Semiparametric Inference
Tri Dao, Govinda M Kamath, Vasilis Syrgkanis +1
A popular approach to model compression is to train an inexpensive student model to mimic the class probabilities of a highly accurate but cumbersome teacher model. Surprisingly, t…