activity
20182026
most citedV-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

39 citations · 158 across the 22 of their papers we have counts for

collaborators
Showing 2019Show all

7 papers · 1 filter

cs.LG20194 cited

Quinoa: a Q-function You Infer Normalized Over Actions

Jonas Degrave, Abbas Abdolmaleki, Jost Tobias Springenberg +2

We present an algorithm for learning an approximate action-value soft Q-function in the relative entropy regularised reinforcement learning setting, for which an optimal improved p…

cs.RO20196 cited

Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models

Arunkumar Byravan, Jost Tobias Springenberg, Abbas Abdolmaleki +6

Humans are masters at quickly learning many complex tasks, relying on an approximate understanding of the dynamics of their environments. In much the same way, we would like our le…

cs.AI201939 cited

V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg +11

Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy…

cs.LG2019

Compositional Transfer in Hierarchical Reinforcement Learning

Markus Wulfmeier, Abbas Abdolmaleki, Roland Hafner +7

The successful application of general reinforcement learning algorithms to real-world robotics applications is often limited by their high data requirements. We introduce Regulariz…

cs.LG2019

Robust Reinforcement Learning for Continuous Control with Model Misspecification

Daniel J. Mankowitz, Nir Levine, Rae Jeong +7

We provide a framework for incorporating robustness -- to perturbations in the transition dynamics which we refer to as model misspecification -- into continuous control Reinforcem…

cs.LG20196 cited

Simultaneously Learning Vision and Feature-based Control Policies for Real-world Ball-in-a-Cup

Devin Schwab, Tobias Springenberg, Murilo F. Martins +7

We present a method for fast training of vision based control policies on real robots. The key idea behind our method is to perform multi-task Reinforcement Learning with auxiliary…