44 citations · 159 across the 33 of their papers we have counts for
5 papers · 1 filter
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy +3
Reinforcement learning from human feedback (RLHF) has emerged as a central tool for language model alignment. We consider online exploration in RLHF, which exploits interactive acc…
The Power of Resets in Online Reinforcement Learning
Zakaria Mhammedi, Dylan J. Foster, Alexander Rakhlin
Simulators are a pervasive tool in reinforcement learning, but most existing algorithms cannot efficiently exploit simulator access -- particularly in high-dimensional domains that…
Online Estimation via Offline Estimation: An Information-Theoretic Framework
Dylan J. Foster, Yanjun Han, Jian Qian +1
The classical theory of statistical estimation aims to estimate a parameter of interest under data generated from a fixed design ("offline estimation"), while the contemporary t…
Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
Zeyu Jia, Alexander Rakhlin, Ayush Sekhari +1
We revisit the problem of offline reinforcement learning with value function realizability but without Bellman completeness. Previous work by Xie and Jiang (2021) and Foster et al.…
On the Performance of Empirical Risk Minimization with Smoothed Data
Adam Block, Alexander Rakhlin, Abhishek Shetty
In order to circumvent statistical and computational hardness results in sequential decision-making, recent work has considered smoothed online learning, where the distribution of…