2 papers
stat.ML2020
To update or not to update? Delayed Nonparametric Bandits with Randomized Allocation
Sakshi Arya, Yuhong Yang
Delayed rewards problem in contextual bandits has been of interest in various practical settings. We study randomized allocation strategies and provide an understanding on how the…
stat.ML2019
Randomized Allocation with Nonparametric Estimation for Contextual Multi-Armed Bandits with Delayed Rewards
Sakshi Arya, Yuhong Yang
We study a multi-armed bandit problem with covariates in a setting where there is a possible delay in observing the rewards. Under some mild assumptions on the probability distribu…