activity
20172026
most citedTwo Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples

42 citations · 268 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

27 papers · 1 filter

cs.LG2026

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails

Ruinan Jin, Yingbin Liang, Shaofeng Zou

Despite Adam demonstrating faster empirical convergence than SGD in many applications, much of the existing theory yields guarantees essentially comparable to those of SGD, leaving…

cs.LG2022

Data Sampling Affects the Complexity of Online SGD over Dependent Data

Shaocong Ma, Ziyi Chen, Yi Zhou +2

Conventional machine learning applications typically assume that data samples are independently and identically distributed (i.i.d.). However, practical scenarios often involve a d…

cs.LG20211 cited

A Unified Off-Policy Evaluation Approach for General Value Function

Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1

General Value Function (GVF) is a powerful tool to represent both the {\em predictive} and {\em retrospective} knowledge in reinforcement learning (RL). In practice, often multiple…

cs.LG2021

Doubly Robust Off-Policy Actor-Critic: Convergence and Optimality

Tengyu Xu, Zhuoran Yang, Zhaoran Wang +1

Designing off-policy reinforcement learning algorithms is typically a very challenging task, because a desirable iteration update often involves an expectation over an on-policy di…

cs.LG20205 cited

Sample Complexity Bounds for Two Timescale Value-based Reinforcement Learning Algorithms

Tengyu Xu, Yingbin Liang

Two timescale stochastic approximation (SA) has been widely used in value-based reinforcement learning algorithms. In the policy evaluation setting, it can model the linear and non…

cs.LG2020

CRPO: A New Approach for Safe Reinforcement Learning with Convergence Guarantee

Tengyu Xu, Yingbin Liang, Guanghui Lan

In safe reinforcement learning (SRL) problems, an agent explores the environment to maximize an expected total reward and meanwhile avoids violation of certain constraints on a num…