activity
20172026
most citedBenchmarking Offline Reinforcement Learning on Real-Robot Hardware

11 citations · 12 across the 7 of their papers we have counts for

collaborators

14 papers

cs.RO2026

LoComposition: Terrain-Adaptive Energy-Efficient Quadruped Locomotion without Gait Priors

Loukas Kordos, Leonard T. Franz, Simon Rappenecker +4

Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation…

cs.LG2026

GASP: Guided Asymmetric Self-Play For Coding LLMs

Swadesh Jana, Cansu Sancaktar, Tomáš Daniš +3

Asymmetric self-play has emerged as a promising paradigm for post-training large language models, where a teacher continually generates questions for a student to solve at the edge…

cs.LG2026

Stochastic Decision Horizons for Constrained Reinforcement Learning

Nikola Milosevic, Leonard Franz, Daniel Haeufle +3

We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint satisfaction, a desirable property…

cs.AI2025

SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World Models

Cansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk +2

Exploration is a cornerstone of reinforcement learning (RL). Intrinsic motivation attempts to decouple exploration from external, task-based rewards. However, established approache…

cs.LG2025

Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints

Pavel Kolev, Marin Vlastelica, Georg Martius

Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, improving robustness to distribution shift w…

cs.RO2023

Learning Diverse Skills for Local Navigation under Multi-constraint Optimality

Jin Cheng, Marin Vlastelica, Pavel Kolev +2

Despite many successful applications of data-driven control in robotics, extracting meaningful diverse behaviors remains a challenge. Typically, task performance needs to be compro…