2 papers
cs.LG2026
Distributional Reinforcement Learning via the Cramér Distance
Vanya Aziz, Ivo Nowak, E. M. T Hendrix
This paper explores the application of the Soft Actor-Critic (SAC) algorithm within a Distributional Reinforcement Learning setting and introduces an implementation of such algorit…
cs.LG2026
StructRL: Recovering Dynamic Programming Structure from Learning Dynamics in Distributional Reinforcement Learning
Ivo Nowak
Reinforcement learning is typically treated as a uniform, data-driven optimization process, where updates are guided by rewards and temporal-difference errors without explicitly ex…