187 citations · 217 across the 13 of their papers we have counts for
15 papers
Discrete Factorial Representations as an Abstraction for Goal Conditioned Reinforcement Learning
Riashat Islam, Hongyu Zang, Anirudh Goyal +6
Goal-conditioned reinforcement learning (RL) is a promising direction for training agents that are capable of solving multiple tasks and reach a diverse set of objectives. How to \…
Non-Markovian policies occupancy measures
Romain Laroche, Remi Tachet des Combes, Jacob Buckman
A central object of study in Reinforcement Learning (RL) is the Markovian policy, in which an agent's actions are chosen from a memoryless probability distribution, conditioned onl…
One-Shot Learning from a Demonstration with Hierarchical Latent Language
Nathaniel Weir, Xingdi Yuan, Marc-Alexandre Côté +5
Humans have the capability, aided by the expressive compositionality of their language, to learn quickly by demonstration. They are able to describe unseen task-performing procedur…
Beyond the Policy Gradient Theorem for Efficient Policy Updates in Actor-Critic Algorithms
Romain Laroche, Remi Tachet
In Reinforcement Learning, the optimal action at a given state is dependent on policy decisions at subsequent states. As a consequence, the learning targets evolve with time and th…
Batched Bandits with Crowd Externalities
Romain Laroche, Othmane Safsafi, Raphael Feraud +1
In Batched Multi-Armed Bandits (BMAB), the policy is not allowed to be updated at each time step. Usually, the setting asserts a maximum number of allowed policy updates and the al…
Dr Jekyll and Mr Hyde: the Strange Case of Off-Policy Policy Updates
Romain Laroche, Remi Tachet
The policy gradient theorem states that the policy should only be updated in states that are visited by the current policy, which leads to insufficient planning in the off-policy s…