3 papers
cs.LG2026
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically…
cs.RO2026
Guided Discovery of New Behaviors using Diffusion Policies
Dian Yu, Sebastian Sanokowski, Majid Khadiv
Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajectory distributions. However,…
cs.RO2026
Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering
Aditya Shirwatkar, Sebastian Sanokowski, Shishir Kolathaya +2
Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent during training. Large-scale off…