11 citations · 14 across the 6 of their papers we have counts for
8 papers
Continuous-in-time Limit for Bayesian Bandits
Yuhua Zhu, Zachary Izzo, Lexing Ying
This paper revisits the bandit problem in the Bayesian setting. The Bayesian approach formulates the bandit problem as an optimization problem, and the goal is to find the optimal…
Operator Shifting for Model-based Policy Evaluation
Xun Tang, Lexing Ying, Yuhua Zhu
In model-based reinforcement learning, the transition matrix and reward vector are often estimated from random samples subject to noise. Even if the estimated model is an unbiased…
Variational Actor-Critic Algorithms
Yuhua Zhu, Lexing Ying
We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variationa…
A Note on Optimization Formulations of Markov Decision Processes
Lexing Ying, Yuhua Zhu
This note summarizes the optimization formulations used in the study of Markov decision processes. We consider both the discounted and undiscounted processes under the standard and…
Why resampling outperforms reweighting for correcting sampling bias with stochastic gradients
Jing An, Lexing Ying, Yuhua Zhu
A data set sampled from a certain population is biased if the subgroups of the population are sampled at proportions that are significantly different from their underlying proporti…
Borrowing From the Future: Addressing Double Sampling in Model-free Control
Yuhua Zhu, Zach Izzo, Lexing Ying
In model-free reinforcement learning, the temporal difference method and its variants become unstable when combined with nonlinear function approximations. Bellman residual minimiz…