22 citations · 26 across the 3 of their papers we have counts for
8 papers
TDM: Trustworthy Decision-Making via Interpretability Enhancement
Daoming Lyu, Fangkai Yang, Hugh Kwon +3
Human-robot interactive decision-making is increasingly becoming ubiquitous, and trust is an influential factor in determining the reliance on autonomy. However, it is not reasonab…
Variance-Reduced Off-Policy Memory-Efficient Policy Search
Daoming Lyu, Qi Qi, Mohammad Ghavamzadeh +3
Off-policy policy optimization is a challenging problem in reinforcement learning (RL). The algorithms designed for this problem often suffer from high variance in their estimators…
GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values
Shangtong Zhang, Bo Liu, Shimon Whiteson
We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. Gra…
Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
Shangtong Zhang, Bo Liu, Hengshuai Yao +1
We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic,…
A Human-Centered Data-Driven Planner-Actor-Critic Architecture via Logic Programming
Daoming Lyu, Fangkai Yang, Bo Liu +1
Recent successes of Reinforcement Learning (RL) allow an agent to learn policies that surpass human experts but suffers from being time-hungry and data-hungry. By contrast, human l…
QUOTA: The Quantile Option Architecture for Reinforcement Learning
Shangtong Zhang, Borislav Mavrin, Linglong Kong +2
In this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL). In QUOTA, decision making…