3 papers
stat.ML2026
Variance-aware Reward Modeling with Anchor Guidance
Shuxing Fang, Ruijian Han, Liangyu Zhang +1
Standard Bradley--Terry (BT) reward models are limited when human preferences are pluralistic. Although soft preference labels preserve disagreement information, BT can only expres…
stat.ML2025
A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
Yang Peng, Kaicheng Jin, Liangyu Zhang +1
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD lea…
stat.ML2025
Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces
Yang Peng, Liangyu Zhang, Zhihua Zhang
Distributional reinforcement learning (DRL) has achieved empirical success in various domains. One core task in DRL is distributional policy evaluation, which involves estimating t…