6 papers
Dominant Arm Identification with Mixing and Recycling Observed Samples
Jonghyun Sim, Wonyoung Kim
We study the problem of identifying the dominant arm in multi-armed bandits, where the objective is to find the action with the highest probability of exceeding the realized reward…
Variance-Adaptive Optimal Algorithm for Reinforcement Learning with Multinomial Logit Function Approximation
Wonyoung Kim, Min-Hwan Oh, Garud Iyengar +1
Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studi…
DiTTO-LLM: Framework for Discovering Topic-based Technology Opportunities via Large Language Model
Wonyoung Kim, Sujeong Seo, Juhyun Lee
Technology opportunities are critical information that serve as a foundation for advancements in technology, industry, and innovation. This paper proposes a framework based on the…
Linear Bandits with Partially Observable Features
Wonyoung Kim, Sungwoo Park, Garud Iyengar +2
We study the linear bandit problem that accounts for partially observable features. Without proper handling, unobserved features can lead to linear regret in the decision horizon $…
Adaptive Data Augmentation for Thompson Sampling
Wonyoung Kim
In linear contextual bandits, the objective is to select actions that maximize cumulative rewards, modeled as a linear function with unknown parameters. Although Thompson Sampling…
Learning the Pareto Front Using Bootstrapped Observation Samples
Wonyoung Kim, Garud Iyengar, Assaf Zeevi
We consider Pareto front identification (PFI) for linear bandits (PFILin), i.e., the goal is to identify a set of arms with undominated mean reward vectors when the mean reward vec…