1 citations · 3 across the 11 of their papers we have counts for
11 papers · 1 filter
Dominant Arm Identification with Mixing and Recycling Observed Samples
Jonghyun Sim, Wonyoung Kim
We study the problem of identifying the dominant arm in multi-armed bandits, where the objective is to find the action with the highest probability of exceeding the realized reward…
Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation
Wonyoung Kim, Min-Hwan Oh, Garud Iyengar +1
Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studi…
Adaptive Data Augmentation for Thompson Sampling
Wonyoung Kim
In linear contextual bandits, the objective is to select actions that maximize cumulative rewards, modeled as a linear function with unknown parameters. Although Thompson Sampling…
Linear Bandits with Partially Observable Features
Wonyoung Kim, Sungwoo Park, Garud Iyengar +2
We study the linear bandit problem that accounts for partially observable features. Without proper handling, unobserved features can lead to linear regret in the decision horizon $…
A Doubly Robust Approach to Sparse Reinforcement Learning
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
We propose a new regret minimization algorithm for episodic sparse linear Markov decision process (SMDP) where the state-transition distribution is a linear function of observed fe…
Learning the Pareto Front Using Bootstrapped Observation Samples
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
We consider Pareto front identification (PFI) for linear bandits (PFILin), i.e., the goal is to identify a set of arms with undominated mean reward vectors when the mean reward vec…