collaborators

5 papers

math.ST2025

On the Variance, Admissibility, and Stability of Empirical Risk Minimization

Gil Kur, Eli Putterman, Alexander Rakhlin

It is well known that Empirical Risk Minimization (ERM) may attain minimax suboptimal rates in terms of the mean squared error (Birgé and Massart, 1993). In this paper, we prove t…

math.PR2025

Rate of convergence of the smoothed empirical Wasserstein distance

Adam Block, Zeyu Jia, Yury Polyanskiy +1

Consider an empirical measure induced by iid samples from a -dimensional -subgaussian distribution and let be the isotropi…

cs.LG2024

Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy +3

Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment. We consider online exploration in RLHF, which exploits interactive acc…

cs.LG2024

The Power of Resets in Online Reinforcement Learning

Zakaria Mhammedi, Dylan J. Foster, Alexander Rakhlin

Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that…

stat.ML2024

Online Estimation via Offline Estimation: An Information-Theoretic Framework

Dylan J. Foster, Yanjun Han, Jian Qian +1

The classical theory of statistical estimation aims to estimate a parameter of interest under data generated from a fixed design ("offline estimation"), while the contemporary t…