45 citations · 182 across the 44 of their papers we have counts for
Showing 2022 · math.OCShow all
2 papers · 2 filters
math.OC2022
Continuous-in-time Limit for Bayesian Bandits
Yuhua Zhu, Zachary Izzo, Lexing Ying
This paper revisits the bandit problem in the Bayesian setting. The Bayesian approach formulates the bandit problem as an optimization problem, and the goal is to find the optimal…
math.OC2022
Accelerating Primal-dual Methods for Regularized Markov Decision Processes
Haoya Li, Hsiang-fu Yu, Lexing Ying +1
Entropy regularized Markov decision processes have been widely used in reinforcement learning. This paper is concerned with the primal-dual formulation of the entropy regularized p…