3 papers
cs.LG2025
Infrequent Exploration in Linear Bandits
Harin Lee, Min-hwan Oh
We study the problem of infrequent exploration in linear bandits, addressing a significant yet overlooked gap between fully adaptive exploratory methods (e.g., UCB and Thompson Sam…
cs.LG2025
Minimax Optimal Reinforcement Learning with Quasi-Optimism
Harin Lee, Min-hwan Oh
In our quest for a reinforcement learning (RL) algorithm that is both practical and provably optimal, we introduce EQO (Exploration via Quasi-Optimism). Unlike existing minimax opt…
stat.ML2024
Improved Regret of Linear Ensemble Sampling
Harin Lee, Min-hwan Oh
In this work, we close the fundamental gap of theory and practice by providing an improved regret bound for linear ensemble sampling. We prove that with an ensemble size logarithmi…