22 citations · 26 across the 3 of their papers we have counts for
10 papers · 1 filter
STOPS: Short-Term-based Volatility-controlled Policy Search and its Global Convergence
Liangliang Xu, Daoming Lyu, Yangchen Pan +2
It remains challenging to deploy existing risk-averse approaches to real-world applications. The reasons are multi-fold, including the lack of global optimality guarantee and the n…
TDM: Trustworthy Decision-Making via Interpretability Enhancement
Daoming Lyu, Fangkai Yang, Hugh Kwon +3
Human-robot interactive decision-making is increasingly becoming ubiquitous, and trust is an influential factor in determining the reliance on autonomy. However, it is not reasonab…
Variance-Reduced Off-Policy Memory-Efficient Policy Search
Daoming Lyu, Qi Qi, Mohammad Ghavamzadeh +3
Off-policy policy optimization is a challenging problem in reinforcement learning (RL). The algorithms designed for this problem often suffer from high variance in their estimators…
Stable and Efficient Policy Evaluation
Daoming Lyu, Bo Liu, Matthieu Geist +3
Policy evaluation algorithms are essential to reinforcement learning due to their ability to predict the performance of a policy. However, there are two long-standing issues lying…
Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning
Shangtong Zhang, Bo Liu, Shimon Whiteson
We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variab…
GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values
Shangtong Zhang, Bo Liu, Shimon Whiteson
We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. Gra…