4 papers · 1 filter
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL
Franki Nguimatsia-Tiofack, Fabian Schramm, Théotime Le Hellard +1
While self-supervised Contrastive Reinforcement Learning (CRL) has shown remarkable depth-scaling capabilities, successfully using networks over 64 layers, scaled CRL still struggl…
SVL: Goal-Conditioned Reinforcement Learning as Survival Learning
Franki Nguimatsia Tiofack, Fabian Schramm, Théotime Le Hellard +1
Standard approaches to goal-conditioned reinforcement learning (GCRL) that rely on temporal-difference learning can be unstable and sample-inefficient due to bootstrapping. While r…
Accelerating trajectory optimization with Sobolev-trained diffusion policies
Théotime Le Hellard, Franki Nguimatsia Tiofack, Quentin Le Lidec +1
Trajectory Optimization (TO) solvers exploit known system dynamics to compute locally optimal trajectories through iterative improvements. A downside is that each new problem insta…
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
Franki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm +2
Offline reinforcement learning often relies on behavior regularization that enforces policies to remain close to the dataset distribution. However, such approaches fail to distingu…