10 papers
Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification
Joongkyu Lee, Min-hwan Oh
We study optimal experimental design for multinomial logit (MNL) bandits, where an agent repeatedly selects a subset of items from a ground set of size and observes single-…
Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent
Joongkyu Lee, Min-hwan Oh
We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-varying parameter. This framew…
Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards
Deokgyu Yoon, Hyungkyu Kang, Joongkyu Lee +4
Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objective…
Block-Sphere Vector Quantization
Heesang Ann, Joongkyu Lee, Min-hwan Oh
Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast retrieval, and compressed inference. Recent rotation-b…
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
Joongkyu Lee, Seouh-won Yi, Min-hwan Oh
We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of theoretical work has emerged-motivated by PbRL'…
Nearly Minimax Optimal Regret for Multinomial Logistic Bandit
Joongkyu Lee, Min-hwan Oh
In this paper, we study the contextual multinomial logit (MNL) bandit problem in which a learning agent sequentially selects an assortment based on contextual information, and user…