4 papers
Dynamics Distillation for Efficient and Transferable Control Learning
Xunjiang Gu, Kashyap Chitta, Mahsa Golchoubian +2
Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, properties that existing simulato…
Relative Entropy Pathwise Policy Optimization
Claas Voelcker, Axel Brunnbauer, Marcel Hussing +6
Score-function based methods for policy learning, such as REINFORCE and PPO, have delivered strong results in game-playing and robotics, yet their high variance often undermines tr…
Calibrated Value-Aware Model Learning with Probabilistic Environment Models
Claas Voelcker, Anastasiia Pedan, Arash Ahmadian +3
The idea of value-aware model learning, that models should produce accurate value estimates, has gained prominence in model-based reinforcement learning. The MuZero loss, which pen…
MAD-TD: Model-Augmented Data stabilizes High Update Ratio RL
Claas A Voelcker, Marcel Hussing, Eric Eaton +2
Building deep reinforcement learning (RL) agents that find a good policy with few samples has proven notoriously challenging. To achieve sample efficiency, recent work has explored…