5 papers
Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Harin Lee, Min-hwan Oh
We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a distributional regret bound as…
Infrequent Exploration in Linear Bandits
Harin Lee, Min-hwan Oh
We study the problem of infrequent exploration in linear bandits, addressing a significant yet overlooked gap between fully adaptive exploratory methods (e.g., UCB and Thompson Sam…
Minimax Optimal Reinforcement Learning with Quasi-Optimism
Harin Lee, Min-hwan Oh
In our quest for a reinforcement learning (RL) algorithm that is both practical and provably optimal, we introduce EQO (Exploration via Quasi-Optimism). Unlike existing minimax opt…
Improved Regret of Linear Ensemble Sampling
Harin Lee, Min-hwan Oh
In this work, we close the fundamental gap of theory and practice by providing an improved regret bound for linear ensemble sampling. We prove that with an ensemble size logarithmi…
Local Anti-Concentration Class: Logarithmic Regret for Greedy Linear Contextual Bandit
Seok-Jin Kim, Min-hwan Oh
We study the performance guarantees of exploration-free greedy algorithms for the linear contextual bandit problem. We introduce a novel condition, named the \textit{Local Anti-Con…