5 papers
LoComposition: Terrain-Adaptive Energy-Efficient Quadruped Locomotion without Gait Priors
Loukas Kordos, Leonard T. Franz, Simon Rappenecker +4
Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation…
Stochastic Decision Horizons for Constrained Reinforcement Learning
Nikola Milosevic, Leonard Franz, Daniel Haeufle +3
We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint satisfaction, a desirable property…
GASP: Guided Asymmetric Self-Play For Coding LLMs
Swadesh Jana, Cansu Sancaktar, Tomáš Daniš +3
Asymmetric self-play has emerged as a promising paradigm for post-training large language models, where a teacher continually generates questions for a student to solve at the edge…
SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models
Cansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk +2
Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approache…
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
Pavel Kolev, Marin Vlastelica, Georg Martius
Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift w…