activity
20182022
most citedTwo Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples

42 citations · 93 across the 11 of their papers we have counts for

collaborators

16 papers

cs.LG20221 cited

Robust Constrained Reinforcement Learning

Yue Wang, Fei Miao, Shaofeng Zou

Constrained reinforcement learning is to maximize the expected reward subject to constraints on utilities/costs. However, the training environment may not be the same as the test o…

cs.LG202211 cited

Policy Gradient Method For Robust Reinforcement Learning

Yue Wang, Shaofeng Zou

This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch. Robust reinfor…

eess.SP202213 cited

Quickest Change Detection in Anonymous Heterogeneous Sensor Networks

Zhongchang Sun, Shaofeng Zou, Ruizhi Zhang +1

The problem of quickest change detection (QCD) in anonymous heterogeneous sensor networks is studied. There are heterogeneous sensors and a fusion center. The sensors are clust…

cs.LG20216 cited

Online Robust Reinforcement Learning with Model Uncertainty

Yue Wang, Shaofeng Zou

Robust reinforcement learning (RL) is to find a policy that optimizes the worst-case performance over an uncertainty set of MDPs. In this paper, we focus on model-free robust RL, w…

math.OC20213 cited

Faster Algorithm and Sharper Analysis for Constrained Markov Decision Process

Tianjiao Li, Ziwei Guan, Shaofeng Zou +3

The problem of constrained Markov decision process (CMDP) is investigated, where an agent aims to maximize the expected accumulated discounted reward subject to multiple constraint…

math.ST20211 cited

Sequential (Quickest) Change Detection: Classical Results and New Directions

Liyan Xie, Shaofeng Zou, Yao Xie +1

Online detection of changes in stochastic systems, referred to as sequential change detection or quickest change detection, is an important research topic in statistics, signal pro…