2 papers
cs.LG2026
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
Jiazheng Sun, Weixin Wang, Pan Xu
We provide a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits and develop corresponding regret bounds for two most common nonlinear contextual ba…
cs.LG2025
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
Weixin Wang, Haoyang Zheng, Guang Lin +2
Most existing approximate Thompson Sampling (TS) algorithms for multi-armed bandits use Stochastic Gradient Langevin Dynamics (SGLD) or its variants in each round to sample from th…