activity
20172021
most citedAgent Modeling as Auxiliary Task for Deep Reinforcement Learning

15 citations · 46 across the 6 of their papers we have counts for

collaborators
Showing 2019Show all

6 papers · 1 filter

cs.LG2019

On Hard Exploration for Reinforcement Learning: a Case Study in Pommerman

Chao Gao, Bilal Kartal, Pablo Hernandez-Leal +1

How to best explore in domains with sparse, delayed, and deceptive rewards is an important open problem for reinforcement learning (RL). This paper considers one such domain, the r…

cs.LG2019

Action Guidance with MCTS for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

Deep reinforcement learning has achieved great successes in recent years, however, one main challenge is the sample inefficiency. In this paper, we focus on how to use action guida…

cs.LG2019

Terminal Prediction as an Auxiliary Task for Deep Reinforcement Learning

Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor

Deep reinforcement learning has achieved great successes in recent years, but there are still open challenges, such as convergence to locally optimal policies and sample inefficien…

cs.MA2019★ 15 cited

Agent Modeling as Auxiliary Task for Deep Reinforcement Learning

Pablo Hernandez-Leal, Bilal Kartal, Matthew E. Taylor

In this paper we explore how actor-critic methods in deep reinforcement learning, in particular Asynchronous Advantage Actor-Critic (A3C), can be extended with agent modeling. Insp…

cs.MA2019★ 14 cited

Skynet: A Top Deep RL Agent in the Inaugural Pommerman Team Competition

Chao Gao, Pablo Hernandez-Leal, Bilal Kartal +1

The Pommerman Team Environment is a recently proposed benchmark which involves a multi-agent domain with challenges such as partial observability, decentralized execution (without…

cs.LG2019★ 5 cited

Safer Deep RL with Shallow MCTS: A Case Study in Pommerman

Bilal Kartal, Pablo Hernandez-Leal, Chao Gao +1

Safe reinforcement learning has many variants and it is still an open research problem. Here, we focus on how to use action guidance by means of a non-expert demonstrator to avoid…