activity
20182022
most citedTDM: Trustworthy Decision-Making via Interpretability Enhancement

22 citations · 26 across the 3 of their papers we have counts for

collaborators

8 papers

cs.LG202122 cited

TDM: Trustworthy Decision-Making via Interpretability Enhancement

Daoming Lyu, Fangkai Yang, Hugh Kwon +3

Human-robot interactive decision-making is increasingly becoming ubiquitous, and trust is an influential factor in determining the reliance on autonomy. However, it is not reasonab…

cs.LG20204 cited

Variance-Reduced Off-Policy Memory-Efficient Policy Search

Daoming Lyu, Qi Qi, Mohammad Ghavamzadeh +3

Off-policy policy optimization is a challenging problem in reinforcement learning (RL). The algorithms designed for this problem often suffer from high variance in their estimators…

cs.LG2020

GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values

Shangtong Zhang, Bo Liu, Shimon Whiteson

We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. Gra…

cs.LG2019

Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation

Shangtong Zhang, Bo Liu, Hengshuai Yao +1

We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic,…

cs.AI2019

A Human-Centered Data-Driven Planner-Actor-Critic Architecture via Logic Programming

Daoming Lyu, Fangkai Yang, Bo Liu +1

Recent successes of Reinforcement Learning (RL) allow an agent to learn policies that surpass human experts but suffers from being time-hungry and data-hungry. By contrast, human l…

cs.LG2018

QUOTA: The Quantile Option Architecture for Reinforcement Learning

Shangtong Zhang, Borislav Mavrin, Linglong Kong +2

In this paper, we propose the Quantile Option Architecture (QUOTA) for exploration based on recent advances in distributional reinforcement learning (RL). In QUOTA, decision making…