collaborators

9 papers

cs.LG2026

Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

Huy Le, Onur Celik, Denis Blessing +6

Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely…

cs.RO2026

Update-Free On-Policy Steering via Verifiers

Maria Attarian, Ian Vyse, Claas Voelcker +7

In recent years, Behavior Cloning (BC) has become one of the most prevalent methods for learning manipulation from human demonstrations. Despite their successes, BC policies are of…

cs.AI2026

Exploiting Local Dynamics Regularity for Reusable Skills in Offline Hierarchical RL

Sarthak Dayal, Abhinav Peri, Carl Qi +4

Hierarchical Reinforcement Learning (HRL) promises to solve long-horizon Reinforcement Learning (RL) tasks more efficiently than non-hierarchical counterparts by discovering and re…

cs.LG2026

Test-Time Graph Search for Goal-Conditioned Reinforcement Learning

Evgenii Opryshko, Junwei Quan, Claas Voelcker +2

Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is t…

cs.LG2026

Behavior-Consistent Deep Reinforcement Learning

Marcel Hussing, Liv G. d'Aliberti, Claas Voelcker +2

Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challenge to deployment in real-world domains. I…

cs.LG2026

Relative Entropy Pathwise Policy Optimization

Claas Voelcker, Axel Brunnbauer, Marcel Hussing +6

Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines tr…