187 citations · 216 across the 5 of their papers we have counts for
8 papers
One-Shot Learning from a Demonstration with Hierarchical Latent Language
Nathaniel Weir, Xingdi Yuan, Marc-Alexandre Côté +5
Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedur…
Shortest-Path Constrained Reinforcement Learning for Sparse Reward Tasks
Sungryull Sohn, Sungtae Lee, Jongwook Choi +3
We propose the k-Shortest-Path (k-SP) constraint: a novel constraint on the agent's trajectory that improves the sample efficiency in sparse-reward MDPs. We show that any optimal p…
The LoCA Regret: A Consistent Metric to Evaluate Model-Based Behavior in Reinforcement Learning
Harm van Seijen, Hadi Nekoei, Evan Racah +1
Deep model-based Reinforcement Learning (RL) has the potential to substantially improve the sample-efficiency of deep RL. While various challenges have long held it back, a number…
Using a Logarithmic Mapping to Enable Lower Discount Factors in Reinforcement Learning
Harm van Seijen, Mehdi Fatemi, Arash Tavakoli
In an effort to better understand the different ways in which the discount factor affects the optimization process in reinforcement learning, we designed a set of experiments to st…
Learning Invariances for Policy Generalization
Remi Tachet, Philip Bachman, Harm van Seijen
While recent progress has spawned very powerful machine learning systems, those agents remain extremely specialized and fail to transfer the knowledge they gain to similar yet unse…
Hybrid Reward Architecture for Reinforcement Learning
Harm van Seijen, Mehdi Fatemi, Joshua Romoff +3
One of the main challenges in reinforcement learning (RL) is generalisation. In typical deep RL methods this is achieved by approximating the optimal value function with a low-dime…