collaborators

13 papers

cs.LG2026

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

Hyungkyu Kang, Byeongchan Kim, Min-hwan Oh

Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However, learning a reliable goal-co…

stat.ML2026

Optimal Design for Multinomial Logit Model with Applications to Best Assortment Identification

Joongkyu Lee, Min-hwan Oh

We study optimal experimental design for multinomial logit (MNL) bandits, where an agent repeatedly selects a subset of items from a ground set of size and observes single-…

stat.ML2026

Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent

Joongkyu Lee, Min-hwan Oh

We study nonstationary generalized linear bandits (GLBs), where the expected reward is modeled through a nonlinear link function with an unknown time-varying parameter. This framew…

cs.LG2026

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

Deokgyu Yoon, Hyungkyu Kang, Joongkyu Lee +4

Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely used PPO surrogate objective…

cs.LG2026

Block-Sphere Vector Quantization

Heesang Ann, Joongkyu Lee, Min-hwan Oh

Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast retrieval, and compressed inference. Recent rotation-b…

cs.LG2026

Peng's Q() for Conservative Value Estimation in Offline Reinforcement Learning

Byeongchan Kim, Min-hwan Oh

We propose a model-free offline multi-step reinforcement learning (RL) algorithm, Conservative Peng's Q() (CPQL). Our algorithm adapts the Peng's Q() (PQL) operator for con…