activity
20182024
most citedV-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control

39 citations · 148 across the 14 of their papers we have counts for

collaborators
Showing cs.LGShow all

16 papers · 1 filter

cs.LG2022

Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach

Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6

Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…

cs.LG20214 cited

Collect & Infer -- a fresh look at data-efficient Reinforcement Learning

Martin Riedmiller, Jost Tobias Springenberg, Roland Hafner +1

This position paper proposes a fresh look at Reinforcement Learning (RL) from the perspective of data-efficiency. Data-efficient RL has gone through three major stages: pure on-lin…

cs.LG2021

Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning

William F. Whitney, Michael Bloesch, Jost Tobias Springenberg +3

Despite the close connection between exploration and sample efficiency, most state of the art reinforcement learning algorithms include no considerations for exploration beyond max…

cs.LG2020

Local Search for Policy Iteration in Continuous Control

Jost Tobias Springenberg, Nicolas Heess, Daniel Mankowitz +10

We present an algorithm for local, regularized, policy improvement in reinforcement learning (RL) that allows us to formulate model-based and model-free variants in a single framew…

cs.LG20203 cited

Simple Sensor Intentions for Exploration

Tim Hertweck, Martin Riedmiller, Michael Bloesch +5

Modern reinforcement learning algorithms can learn solutions to increasingly difficult control problems while at the same time reduce the amount of prior knowledge needed for their…

cs.LG2020

Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning

Noah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp +6

Off-policy reinforcement learning algorithms promise to be applicable in settings where only a fixed data-set (batch) of environment interactions is available and no new experience…