4 papers
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
Franki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm +2
Offline reinforcement learning often relies on behavior regularization that enforces policies to remain close to the dataset distribution. However, such approaches fail to distingu…
First-order Sobolev Reinforcement Learning
Fabian Schramm, Nicolas Perrin-Gilbert, Justin Carpentier
We propose a refinement of temporal-difference learning that enforces first-order Bellman consistency: the learned value function is trained to match not only the Bellman targets i…
Reference-Free Sampling-Based Model Predictive Control
Fabian Schramm, Pierre Fabre, Nicolas Perrin-Gilbert +1
We present a sampling-based model predictive control (MPC) framework that enables emergent locomotion without relying on handcrafted gait patterns or predefined contact sequences.…
End-to-End and Highly-Efficient Differentiable Simulation for Robotics
Quentin Le Lidec, Louis Montaut, Yann de Mont-Marin +2
Over the past few years, robotics simulators have largely improved in efficiency and scalability, enabling them to generate years of simulated data in a few hours. Yet, efficiently…