42 citations · 111 across the 21 of their papers we have counts for
4 papers · 1 filter
Two Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples
Tengyu Xu, Shaofeng Zou, Yingbin Liang
Gradient-based temporal difference (GTD) algorithms are widely used in off-policy learning scenarios. Among them, the two time-scale TD with gradient correction (TDC) algorithm has…
Finite-Sample Analysis for SARSA with Linear Function Approximation
Shaofeng Zou, Tengyu Xu, Yingbin Liang
SARSA is an on-policy algorithm to learn a Markov decision process policy in reinforcement learning. We investigate the SARSA algorithm with linear function approximation under the…
Information-Theoretic Understanding of Population Risk Improvement with Model Compression
Yuheng Bu, Weihao Gao, Shaofeng Zou +1
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in…
Tightening Mutual Information Based Bounds on Generalization Error
Yuheng Bu, Shaofeng Zou, Venugopal V. Veeravalli
An information-theoretic upper bound on the generalization error of supervised learning algorithms is derived. The bound is constructed in terms of the mutual information between e…