4 papers
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
Joongkyu Lee, Seouh-won Yi, Min-hwan Oh
We study online preference-based reinforcement learning (PbRL) with the goal of improving sample efficiency. While a growing body of theoretical work has emerged-motivated by PbRL'…
Combinatorial Reinforcement Learning with Preference Feedback
Joongkyu Lee, Min-hwan Oh
In this paper, we consider combinatorial reinforcement learning with preference feedback, where a learning agent sequentially offers an action--an assortment of multiple items to--…
Improved Online Confidence Bounds for Multinomial Logistic Bandits
Joongkyu Lee, Min-hwan Oh
In this paper, we propose an improved online confidence bound for multinomial logistic (MNL) models and apply this result to MNL bandits, achieving variance-dependent optimal regre…
Demystifying Linear MDPs and Novel Dynamics Aggregation Framework
Joongkyu Lee, Min-hwan Oh
In this work, we prove that, in linear MDPs, the feature dimension is lower bounded by in order to aptly represent transition probabilities, where is the size of the…