3 papers
cs.LG2026
Scalable Maximum Entropy Reinforcement Learning for Diffusion Policies via Adjoint Matching
Serge Thilges, Onur Celik, Denis Blessing +2
Diffusion policies have recently emerged as a powerful paradigm for representing complex action distributions in reinforcement learning (RL). However, their application to online R…
cs.LG2026
PAWS: Preference Learning with Advantage-Weighted Segments
Aleksandar Taranovic, Onur Celik, Niklas Freymuth +6
Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations. Existing methods…
cs.LG2026
TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
Philipp Becker, Niklas Freymuth, Serge Thilges +2
Reinforcement Learning (RL) with PPO-like clip objectives has become the standard choice for reward-based fine-tuning of large language models (LLMs). Although recent work has expl…