6 papers
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Harin Lee, Min-hwan Oh
We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a distributional regret bound as…
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
Harin Lee, Kevin Jamieson
We study reinforcement learning with delayed state observation, where the agent observes the current state after some random number of time steps. We propose an algorithm that comb…
Infrequent Exploration in Linear Bandits
Harin Lee, Min-hwan Oh
We study the problem of infrequent exploration in linear bandits, addressing a significant yet overlooked gap between fully adaptive exploratory methods (e.g., UCB and Thompson Sam…
Minimax Optimal Reinforcement Learning with Quasi-Optimism
Harin Lee, Min-hwan Oh
In our quest for a reinforcement learning (RL) algorithm that is both practical and provably optimal, we introduce EQO (Exploration via Quasi-Optimism). Unlike existing minimax opt…
Improved Regret of Linear Ensemble Sampling
Harin Lee, Min-hwan Oh
In this work, we close the fundamental gap of theory and practice by providing an improved regret bound for linear ensemble sampling. We prove that with an ensemble size logarithmi…
Lasso Bandit with Compatibility Condition on Optimal Arm
Harin Lee, Taehyun Hwang, Min-hwan Oh
We consider a stochastic sparse linear bandit problem where only a sparse subset of context features affects the expected reward function, i.e., the unknown reward parameter has a…