2 papers
cs.LG2026
Multivariate Distributional Reinforcement Learning Using Sliced Divergences
Baptiste Debes, Tinne Tuytelaars
Distributional reinforcement learning (DRL) models the full return distribution rather than expectations, but extending it to multivariate settings remains challenging. Many common…
cs.LG2026
Distributional value gradients for stochastic environments
Baptiste Debes, Tinne Tuytelaars
Gradient-regularized value learning methods improve sample efficiency by leveraging learned models of transition dynamics and rewards to estimate return gradients. However, existin…