10 citations · 24 across the 10 of their papers we have counts for
Showing 2021 · cs.LGShow all
3 papers · 2 filters
cs.LG2021★ 2 cited
Debiasing Samples from Online Learning Using Bootstrap
Ningyuan Chen, Xuefeng Gao, Yi Xiong
It has been recently shown in the literature that the sample averages from online learning experiments are biased when used to estimate the mean reward. To correct the bias, off-po…
cs.LG2021★ 1 cited
Sublinear Regret for Learning POMDPs
Yi Xiong, Ningyuan Chen, Xuefeng Gao +1
We study the model-based undiscounted reinforcement learning for partially observable Markov decision processes (POMDPs). The oracle we consider is the optimal policy of the POMDP…
cs.LG2021★ 1 cited
Multi-armed Bandit Requiring Monotone Arm Sequences
Ningyuan Chen
In many online learning or multi-armed bandit problems, the taken actions or pulled arms are ordinal and required to be monotone over time. Examples include dynamic pricing, in whi…