18 citations · 19 across the 11 of their papers we have counts for
12 papers · 1 filter
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL
Huy Le, Onur Celik, Denis Blessing +6
Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely…
Behavior-Consistent Deep Reinforcement Learning
Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker +2
Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. I…
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
Evgenii Opryshko, Junwei Quan, Claas Voelcker +2
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is t…
Relative Entropy Pathwise Policy Optimization
Claas Voelcker, Axel Brunnbauer, Marcel Hussing +6
Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines tr…
Calibrated Value-Aware Model Learning with Probabilistic Environment Models
Claas Voelcker, Anastasiia Pedan, Arash Ahmadian +3
The idea of value-aware model learning, that models should produce accurate value estimates, has gained prominence in model-based reinforcement learning. The MuZero loss, which pen…
Temporal-Difference Learning Using Distributed Error Signals
Jonas Guan, Shon Eduard Verch, Claas Voelcker +3
A computational problem in biological reward-based learning is how credit assignment is performed in the nucleus accumbens (NAc). Much research suggests that NAc dopamine encodes t…