Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
Franki Nguimatsia Tiofack, Théotime Le Hellard, Fabian Schramm +2
Offline reinforcement learning often relies on behavior regularization that enforces policies to remain close to the dataset distribution. However, such approaches fail to distingu…
cs.LG2025
First-order Sobolev Reinforcement Learning
Fabian Schramm, Nicolas Perrin-Gilbert, Justin Carpentier
We propose a refinement of temporal-difference learning that enforces first-order Bellman consistency: the learned value function is trained to match not only the Bellman targets i…