activity
20212026
collaborators
Showing stat.MLShow all

6 papers · 1 filter

stat.ML2026

A Finite-Sample Analysis of Quantile Temporal-Difference Learning

Zijie Cheng, Xiang Li, Yang Peng +1

Quantile temporal-difference learning (QTD) is an effective method for learning return distributions through quantile approximation, yet its finite-time behavior remains poorly und…

stat.ML2026

Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

Zijie Cheng, Yang Peng, Zhihua Zhang

In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generati…

stat.ML2026

Statistical Efficiency and Inference of Quantile Distributional Reinforcement Learning

Zijie Cheng, Yang Peng, Zhihua Zhang

In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency. We focus on distributional policy evaluation, whose goa…

stat.ML2024

Federated Control in Markov Decision Processes

Hao Jin, Yang Peng, Liangyu Zhang +1

We study problems of federated control in Markov Decision Processes. To solve an MDP with large state space, multiple learning agents are introduced to collaboratively learn its op…

stat.ML2024

Statistical Efficiency of Distributional Temporal Difference Learning and Freedman's Inequality in Hilbert Spaces

Yang Peng, Liangyu Zhang, Zhihua Zhang

Distributional reinforcement learning (DRL) has achieved empirical success in various domains. One core task in DRL is distributional policy evaluation, which involves estimating t…

stat.ML2021

Towards Theoretical Understandings of Robust Markov Decision Processes: Sample Complexity and Asymptotics

Wenhao Yang, Liangyu Zhang, Zhihua Zhang

In this paper, we study the non-asymptotic and asymptotic performances of the optimal robust policy and value function of robust Markov Decision Processes(MDPs), where the optimal…