2 papers
cs.LG2026
Distributional Soft Bellman Operator under the Cramér Geometry
Keru Wang, Yixin Deng, Yao Lyu +2
Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the pol…
cs.LG2026
A Spectral Revisit of the Distributional Bellman Operator under the Cramér Metric
Keru Wang, Yixin Deng, Yao Lyu +2
Distributional reinforcement learning (DRL) studies the evolution of full return distributions under Bellman updates rather than focusing on expected values. A classical result is…