3 citations · 4 across the 3 of their papers we have counts for
4 papers
A Note on Target Q-learning For Solving Finite MDPs with A Generative Oracle
Ziniu Li, Tian Xu, Yang Yu
Q-learning with function approximation could diverge in the off-policy setting and the target network is a powerful technique to address this issue. In this manuscript, we examine…
Rethinking ValueDice: Does It Really Improve Performance?
Ziniu Li, Tian Xu, Yang Yu +1
Since the introduction of GAIL, adversarial imitation learning (AIL) methods attract lots of research interests. Among these methods, ValueDice has achieved significant improvement…
Error Bounds of Imitating Policies and Environments
Tian Xu, Ziniu Li, Yang Yu
Imitation learning trains a policy by mimicking expert demonstrations. Various imitation methods were proposed and empirically evaluated, meanwhile, their theoretical understanding…
On Value Discrepancy of Imitation Learning
Tian Xu, Ziniu Li, Yang Yu
Imitation learning trains a policy from expert demonstrations. Imitation learning approaches have been designed from various principles, such as behavioral cloning via supervised l…