activity
20152026
most citedBenchmarking Batch Deep Reinforcement Learning Algorithms

159 citations · 755 across the 46 of their papers we have counts for

collaborators
Showing 2020Show all

17 papers · 1 filter

cs.LG2020★ 3 cited

Non-Stationary Latent Bandits

Joey Hong, Branislav Kveton, Manzil Zaheer +4

Users of recommender systems often behave in a non-stationary fashion, due to their evolving preferences and tastes over time. In this work, we propose a practical approach for fas…

cs.LG2020

Soft-Robust Algorithms for Batch Reinforcement Learning

Elita A. Lobo, Mohammad Ghavamzadeh, Marek Petrik

In reinforcement learning, robust policies for high-stakes decision-making problems with limited data are usually computed by optimizing the percentile criterion, which minimizes t…

cs.LG2020

A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges

Moloud Abdar, Farhad Pourpanah, Sadiq Hussain +9

Uncertainty quantification (UQ) plays a pivotal role in reduction of uncertainties during both optimization and decision making processes. It can be applied to solve a variety of r…

cs.LG2020★ 4 cited

Variance-Reduced Off-Policy Memory-Efficient Policy Search

Daoming Lyu, Qi Qi, Mohammad Ghavamzadeh +3

Off-policy policy optimization is a challenging problem in reinforcement learning (RL). The algorithms designed for this problem often suffer from high variance in their estimators…

cs.LG2020★ 101 cited

Finite-Sample Analysis of Proximal Gradient TD Algorithms

Bo Liu, Ji Liu, Mohammad Ghavamzadeh +2

In this paper, we analyze the convergence rate of the gradient temporal difference learning (GTD) family of algorithms. Previous analyses of this class of algorithms use ODE techni…

cs.LG2020★ 8 cited

Control-Aware Representations for Model-based Reinforcement Learning

Brandon Cui, Yinlam Chow, Mohammad Ghavamzadeh

A major challenge in modern reinforcement learning (RL) is efficient control of dynamical systems from high-dimensional sensory observations. Learning controllable embedding (LCE)…