activity
20172023
most citedRecurrent Kalman Networks: Factorized Inference in High-Dimensional Deep Feature Spaces

29 citations · 84 across the 21 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG20225 cited

Deep Black-Box Reinforcement Learning with Movement Primitives

Fabian Otto, Onur Celik, Hongyi Zhou +3

\Episode-based reinforcement learning (ERL) algorithms treat reinforcement learning (RL) as a black-box optimization problem where we learn to select a parameter vector of a contro…

cs.LG20223 cited

On Uncertainty in Deep State Space Models for Model-Based Reinforcement Learning

Philipp Becker, Gerhard Neumann

Improved state space models, such as Recurrent State Space Models (RSSMs), are a key factor behind recent advances in model-based reinforcement learning (RL). Yet, despite their em…

cs.LG20223 cited

Specializing Versatile Skill Libraries using Local Mixture of Experts

Onur Celik, Dongzhuoran Zhou, Ge Li +2

A long-cherished vision in robotics is to equip robots with skills that match the versatility and precision of humans. For example, when playing table tennis, a robot should be cap…

cs.LG20216 cited

Differentiable Trust Region Layers for Deep Reinforcement Learning

Fabian Otto, Philipp Becker, Ngo Anh Vien +2

Trust region methods are a popular tool in reinforcement learning as they yield robust policy updates in continuous and discrete action spaces. However, enforcing such trust region…

cs.LG20206 cited

Non-Adversarial Imitation Learning and its Connections to Adversarial Methods

Oleg Arenz, Gerhard Neumann

Many modern methods for imitation learning and inverse reinforcement learning, such as GAIL or AIRL, are based on an adversarial formulation. These methods apply GANs to match the…

cs.LG20206 cited

Expected Information Maximization: Using the I-Projection for Mixture Density Estimation

Philipp Becker, Oleg Arenz, Gerhard Neumann

Modelling highly multi-modal data is a challenging problem in machine learning. Most algorithms are based on maximizing the likelihood, which corresponds to the M(oment)-projection…