activity
20182026
most citedV-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

39 citations · 159 across the 23 of their papers we have counts for

collaborators
Showing 2020Show all

6 papers · 1 filter

stat.ML202012 cited

Training Generative Adversarial Networks by Solving Ordinary Differential Equations

Chongli Qin, Yan Wu, Jost Tobias Springenberg +4

The instability of Generative Adversarial Network (GAN) training has frequently been attributed to gradient descent. Consequently, recent methods have aimed to tailor the models an…

cs.LG2020

Local Search for Policy Iteration in Continuous Control

Jost Tobias Springenberg, Nicolas Heess, Daniel Mankowitz +10

We present an algorithm for local, regularized, policy improvement in reinforcement learning (RL) that allows us to formulate model-based and model-free variants in a single framew…

cs.RO2020

Learning Dexterous Manipulation from Suboptimal Experts

Rae Jeong, Jost Tobias Springenberg, Jackie Kay +5

Learning dexterous manipulation in high-dimensional state-action spaces is an important open challenge with exploration presenting a major bottleneck. Although in many cases the le…

cs.LG20203 cited

Simple Sensor Intentions for Exploration

Tim Hertweck, Martin Riedmiller, Michael Bloesch +5

Modern reinforcement learning algorithms can learn solutions to increasingly difficult control problems while at the same time reduce the amount of prior knowledge needed for their…

cs.LG2020

Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning

Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp +6

Off-policy reinforcement learning algorithms promise to be applicable in settings where only a fixed data-set (batch) of environment interactions is available and no new experience…

cs.LG202027 cited

Continuous-Discrete Reinforcement Learning for Hybrid Control in Robotics

Michael Neunert, Abbas Abdolmaleki, Markus Wulfmeier +7

Many real-world control problems involve both discrete decision variables - such as the choice of control modes, gear switching or digital outputs - as well as continuous decision…