4 papers · 1 filter
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Harin Lee, Min-hwan Oh
We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a distributional regret bound as…
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
Harin Lee, Kevin Jamieson
We study reinforcement learning with delayed state observation, where the agent observes the current state after some random number of time steps. We propose an algorithm that comb…
Infrequent Exploration in Linear Bandits
Harin Lee, Min-hwan Oh
We study the problem of infrequent exploration in linear bandits, addressing a significant yet overlooked gap between fully adaptive exploratory methods (e.g., UCB and Thompson Sam…
Minimax Optimal Reinforcement Learning with Quasi-Optimism
Harin Lee, Min-hwan Oh
In our quest for a reinforcement learning (RL) algorithm that is both practical and provably optimal, we introduce EQO (Exploration via Quasi-Optimism). Unlike existing minimax opt…