activity
20172022
most citedHybrid Reward Architecture for Reinforcement Learning

187 citations · 217 across the 13 of their papers we have counts for

collaborators

15 papers

cs.LG20224 cited

Discrete Factorial Representations as an Abstraction for Goal Conditioned Reinforcement Learning

Riashat Islam, Hongyu Zang, Anirudh Goyal +6

Goal-conditioned reinforcement learning (RL) is a promising direction for training agents that are capable of solving multiple tasks and reach a diverse set of objectives. How to \…

cs.LG2022

Non-Markovian policies occupancy measures

Romain Laroche, Remi Tachet des Combes, Jacob Buckman

A central object of study in Reinforcement Learning (RL) is the Markovian policy, in which an agent's actions are chosen from a memoryless probability distribution, conditioned onl…

cs.CL20222 cited

One-Shot Learning from a Demonstration with Hierarchical Latent Language

Nathaniel Weir, Xingdi Yuan, Marc-Alexandre Côté +5

Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedur…

cs.LG2022

Beyond the Policy Gradient Theorem for Efficient Policy Updates in Actor-Critic Algorithms

Romain Laroche, Remi Tachet

In Reinforcement Learning, the optimal action at a given state is dependent on policy decisions at subsequent states. As a consequence, the learning targets evolve with time and th…

cs.LG2021

Batched Bandits with Crowd Externalities

Romain Laroche, Othmane Safsafi, Raphael Feraud +1

In Batched Multi-Armed Bandits (BMAB), the policy is not allowed to be updated at each time step. Usually, the setting asserts a maximum number of allowed policy updates and the al…

cs.LG20211 cited

Dr Jekyll and Mr Hyde: the Strange Case of Off-Policy Policy Updates

Romain Laroche, Remi Tachet

The policy gradient theorem states that the policy should only be updated in states that are visited by the current policy, which leads to insufficient planning in the off-policy s…