8 papers
Modification-Considering Value Learning for Reward Hacking Mitigation in RL
Evgenii Opryshko, Umangi Jain, Igor Gilitschenski
Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended objective, a failure mode known as reward hacki…
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
Yifan Ruan, Chenyang Cao, Andreas Burger +7
Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning (RL) remains difficult. Effective policy…
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Fangfu Liu, Kai He, Tianchang Shen +7
World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many gen…
Test-Time Graph Search for Goal-Conditioned Reinforcement Learning
Evgenii Opryshko, Junwei Quan, Claas Voelcker +2
Offline goal-conditioned reinforcement learning (GCRL) often struggles with long-horizon tasks, where errors in value estimation accumulate and produce unreliable policies. It is t…
Dynamics Distillation for Efficient and Transferable Control Learning
Xunjiang Gu, Kashyap Chitta, Mahsa Golchoubian +2
Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, properties that existing simulato…
Relative Entropy Pathwise Policy Optimization
Claas Voelcker, Axel Brunnbauer, Marcel Hussing +6
Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines tr…