11 citations · 11 across the 2 of their papers we have counts for
3 papers
Improved Exploration through Latent Trajectory Optimization in Deep Deterministic Policy Gradient
Kevin Sebastian Luck, Mel Vecerik, Simon Stepputtis +2
Model-free reinforcement learning algorithms such as Deep Deterministic Policy Gradient (DDPG) often require additional exploration strategies, especially if the actor is of determ…
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov +13
We present a framework for data-driven robotics that makes use of a large dataset of recorded robot experience and scales to several tasks using learned reward functions. We show h…
Generative predecessor models for sample-efficient imitation learning
Yannick Schroecker, Mel Vecerik, Jonathan Scholz
We propose Generative Predecessor Models for Imitation Learning (GPRIL), a novel imitation learning algorithm that matches the state-action distribution to the distribution observe…