activity
20192022
most citedWhy resampling outperforms reweighting for correcting sampling bias with stochastic gradients

11 citations · 14 across the 6 of their papers we have counts for

collaborators

8 papers

math.OC2022

Continuous-in-time Limit for Bayesian Bandits

Yuhua Zhu, Zachary Izzo, Lexing Ying

This paper revisits the bandit problem in the Bayesian setting. The Bayesian approach formulates the bandit problem as an optimization problem, and the goal is to find the optimal…

cs.LG2021

Operator Shifting for Model-based Policy Evaluation

Xun Tang, Lexing Ying, Yuhua Zhu

In model-based reinforcement learning, the transition matrix and reward vector are often estimated from random samples subject to noise. Even if the estimated model is an unbiased…

cs.LG2021

Variational Actor-Critic Algorithms

Yuhua Zhu, Lexing Ying

We introduce a class of variational actor-critic algorithms based on a variational formulation over both the value function and the policy. The objective function of the variationa…

math.OC2020

A Note on Optimization Formulations of Markov Decision Processes

Lexing Ying, Yuhua Zhu

This note summarizes the optimization formulations used in the study of Markov decision processes. We consider both the discounted and undiscounted processes under the standard and…

cs.LG2020★ 11 cited

Why resampling outperforms reweighting for correcting sampling bias with stochastic gradients

Jing An, Lexing Ying, Yuhua Zhu

A data set sampled from a certain population is biased if the subgroups of the population are sampled at proportions that are significantly different from their underlying proporti…

math.OC2020★ 3 cited

Borrowing From the Future: Addressing Double Sampling in Model-free Control

Yuhua Zhu, Zach Izzo, Lexing Ying

In model-free reinforcement learning, the temporal difference method and its variants become unstable when combined with nonlinear function approximations. Bellman residual minimiz…