2 citations · 3 across the 2 of their papers we have counts for
3 papers
math.OC2022★ 1 cited
A unified algorithm framework for mean-variance optimization in discounted Markov decision processes
Shuai Ma, Xiaoteng Ma, Li Xia
This paper studies the risk-averse mean-variance optimization in infinite-horizon discounted Markov decision processes (MDPs). The involved variance metric concerns reward variabil…
cs.LG2021★ 2 cited
Average-Reward Reinforcement Learning with Trust Region Methods
Xiaoteng Ma, Xiaohang Tang, Li Xia +2
Most of reinforcement learning algorithms optimize the discounted criterion which is beneficial to accelerate the convergence and reduce the variance of estimates. Although the dis…
cs.LG2020
Wasserstein Distance guided Adversarial Imitation Learning with Reward Shape Exploration
Ming Zhang, Yawei Wang, Xiaoteng Ma +4
The generative adversarial imitation learning (GAIL) has provided an adversarial learning framework for imitating expert policy from demonstrations in high-dimensional continuous t…