activity
20132022
most citedHybrid Reward Architecture for Reinforcement Learning

187 citations · 216 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL20222 cited

One-Shot Learning from a Demonstration with Hierarchical Latent Language

Nathaniel Weir, Xingdi Yuan, Marc-Alexandre Côté +5

Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedur…

cs.LG20211 cited

Shortest-Path Constrained Reinforcement Learning for Sparse Reward Tasks

Sungryull Sohn, Sungtae Lee, Jongwook Choi +3

We propose the k-Shortest-Path (k-SP) constraint: a novel constraint on the agent's trajectory that improves the sample efficiency in sparse-reward MDPs. We show that any optimal p…

cs.LG2020

The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning

Harm van Seijen, Hadi Nekoei, Evan Racah +1

Deep model-based Reinforcement Learning (RL) has the potential to substantially improve the sample-efficiency of deep RL. While various challenges have long held it back, a number…

cs.LG2019

Using a Logarithmic Mapping to Enable Lower Discount Factors in Reinforcement Learning

Harm van Seijen, Mehdi Fatemi, Arash Tavakoli

In an effort to better understand the different ways in which the discount factor affects the optimization process in reinforcement learning, we designed a set of experiments to st…

cs.LG2018

Learning Invariances for Policy Generalization

Remi Tachet, Philip Bachman, Harm van Seijen

While recent progress has spawned very powerful machine learning systems, those agents remain extremely specialized and fail to transfer the knowledge they gain to similar yet unse…

cs.LG2017187 cited

Hybrid Reward Architecture for Reinforcement Learning

Harm van Seijen, Mehdi Fatemi, Joshua Romoff +3

One of the main challenges in reinforcement learning (RL) is generalisation. In typical deep RL methods this is achieved by approximating the optimal value function with a low-dime…