activity
20172024
most citedV-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

39 citations · 82 across the 7 of their papers we have counts for

collaborators

10 papers

cs.RO20242 cited

Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning

Dhruva Tirumala, Markus Wulfmeier, Ben Moran +13

We apply multi-agent deep reinforcement learning (RL) to train end-to-end robot soccer policies with fully onboard computation and sensing via egocentric RGB vision. This setting r…

cs.LG20231 cited

Replay across Experiments: A Natural Extension of Off-Policy RL

Dhruva Tirumala, Thomas Lampe, Jose Enrique Chen +9

Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to exte…

cs.LG20221 cited

MO2: Model-Based Offline Options

Sasha Salter, Markus Wulfmeier, Dhruva Tirumala +4

The ability to discover useful behaviours from past experience and transfer them to new tasks is considered a core component of natural embodied intelligence. Inspired by neuroscie…

cs.AI2021

Pick Your Battles: Interaction Graphs as Population-Level Objectives for Strategic Diversity

Marta Garnelo, Wojciech Marian Czarnecki, Siqi Liu +5

Strategic diversity is often essential in games: in multi-player games, for example, evaluating a player against a diverse set of strategies will yield a more accurate estimate of…

cs.AI202014 cited

Behavior Priors for Efficient Reinforcement Learning

Dhruva Tirumala, Alexandre Galashov, Hyeonwoo Noh +8

As we deploy reinforcement learning agents to solve increasingly challenging problems, methods that allow us to inject prior knowledge about the structure of the world and effectiv…

cs.AI201939 cited

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg +11

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy…