5 papers
On the Variance, Admissibility, and Stability of Empirical Risk Minimization
Gil Kur, Eli Putterman, Alexander Rakhlin
It is well known that Empirical Risk Minimization (ERM) may attain minimax suboptimal rates in terms of the mean squared error (Birgé and Massart, 1993). In this paper, we prove t…
Rate of convergence of the smoothed empirical Wasserstein distance
Adam Block, Zeyu Jia, Yury Polyanskiy +1
Consider an empirical measure induced by iid samples from a -dimensional -subgaussian distribution and let be the isotropi…
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy +3
Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment. We consider online exploration in RLHF, which exploits interactive acc…
The Power of Resets in Online Reinforcement Learning
Zakaria Mhammedi, Dylan J. Foster, Alexander Rakhlin
Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that…
Online Estimation via Offline Estimation: An Information-Theoretic Framework
Dylan J. Foster, Yanjun Han, Jian Qian +1
The classical theory of statistical estimation aims to estimate a parameter of interest under data generated from a fixed design ("offline estimation"), while the contemporary t…