activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching

Serge Thilges, Onur Celik, Denis Blessing +2

Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online R…

cs.LG2026

Trust-Region Diffusion Policies for Massively Parallel On-Policy RL

Huy Le, Onur Celik, Denis Blessing +6

Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; however, most existing approaches still rely…

cs.LG2026

PAWS: Preference Learning with Advantage-Weighted Segments

Aleksandar Taranovic, Onur Celik, Niklas Freymuth +6

Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods…

cs.LG2026

Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns

Dong Tian, Onur Celik, Gerhard Neumann

We introduce a sequence-conditioned critic for Soft Actor-Critic (SAC) that models trajectory context with a lightweight Transformer and trains on aggregated -step targets. Unli…

cs.LG2026

SEAR: Sample Efficient Action Chunking Reinforcement Learning

C. F. Maximilian Nagy, Onur Celik, Emiliyan Gospodinov +4

Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended ac…

cs.LG2025

DIME:Diffusion-Based Maximum Entropy Reinforcement Learning

Onur Celik, Zechu Li, Denis Blessing +5

Maximum entropy reinforcement learning (MaxEnt-RL) has become the standard approach to RL due to its beneficial exploration properties. Traditionally, policies are parameterized us…