activity
20182021
most citedAgent Modeling as Auxiliary Task for Deep Reinforcement Learning

15 citations · 46 across the 6 of their papers we have counts for

collaborators

12 papers

cs.LG20211 cited

Robust Risk-Sensitive Reinforcement Learning Agents for Trading Markets

Yue Gao, Kry Yik Chau Lui, Pablo Hernandez-Leal

Trading markets represent a real-world financial application to deploy reinforcement learning agents, however, they carry hard fundamental challenges such as high variance and cost…

cs.LG202011 cited

CDT: Cascading Decision Trees for Explainable Reinforcement Learning

Zihan Ding, Pablo Hernandez-Leal, Gavin Weiguang Ding +2

Deep Reinforcement Learning (DRL) has recently achieved significant advances in various domains. However, explaining the policy of RL agents still remains an open problem due to se…

cs.LG2020

Work in Progress: Temporally Extended Auxiliary Tasks

Craig Sherstan, Bilal Kartal, Pablo Hernandez-Leal +1

Predictive auxiliary tasks have been shown to improve performance in numerous reinforcement learning works, however, this effect is still not well understood. The primary purpose o…

cs.LG2019

On Hard Exploration for Reinforcement Learning: a Case Study in Pommerman

Chao Gao, Bilal Kartal, Pablo Hernandez-Leal +1

How to best explore in domains with sparse, delayed, and deceptive rewards is an important open problem for reinforcement learning (RL). This paper considers one such domain, the r…

cs.LG2019

Action Guidance with MCTS for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

Deep reinforcement learning has achieved great successes in recent years, however, one main challenge is the sample inefficiency. In this paper, we focus on how to use action guida…

cs.LG2019

Terminal Prediction as an Auxiliary Task for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

Deep reinforcement learning has achieved great successes in recent years, but there are still open challenges, such as convergence to locally optimal policies and sample inefficien…