2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 2 cited
Quantum Multi-Armed Bandits and Stochastic Linear Bandits Enjoy Logarithmic Regrets
Zongqi Wan, Zhijie Zhang, Tongyang Li +2
Multi-arm bandit (MAB) and stochastic linear bandit (SLB) are important models in reinforcement learning, and it is well-known that classical algorithms for bandits with time horiz…
cs.LG2022
Bounded Memory Adversarial Bandits with Composite Anonymous Delayed Feedback
Zongqi Wan, Xiaoming Sun, Jialin Zhang
We study the adversarial bandit problem with composite anonymous delayed feedback. In this setting, losses of an action are split into components, spreading over consecutive ro…