collaborators

7 papers

cs.LG2026

Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching

Serge Thilges, Onur Celik, Denis Blessing +2

Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online R…

cs.LG2026

Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

Huy Le, Onur Celik, Denis Blessing +6

Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely…

cs.LG2026

PAWS: Preference Learning with Advantage-Weighted Segments

Aleksandar Taranovic, Onur Celik, Niklas Freymuth +6

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods…

cs.LG2026

Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns

Dong Tian, Onur Celik, Gerhard Neumann

We introduce a sequence-conditioned critic for Soft Actor-Critic (SAC) that models trajectory context with a lightweight Transformer and trains on aggregated -step targets. Unli…

cs.LG2026

SEAR: Sample Efficient Action Chunking Reinforcement Learning

C. F. Maximilian Nagy, Onur Celik, Emiliyan Gospodinov +4

Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended ac…

cs.RO2026

Scaffolding Dexterous Manipulation with Vision-Language Models

Vincent de Bakker, Joey Hejna, Tyler Ga Wei Lum +6

Dexterous robotic hands are essential for performing complex manipulation tasks, yet remain difficult to train due to the challenges of demonstration collection and high-dimensiona…