2 papers
cs.LG2026
Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning
Viktor Veselý, Aleksandar Todorov, Erwan Escudie +1
Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood. We identify a…
cs.LG2025
On The Presence of Double-Descent in Deep Reinforcement Learning
Viktor Veselý, Aleksandar Todorov, Matthia Sabatelli
The double descent (DD) paradox, where over-parameterized models see generalization improve past the interpolation point, remains largely unexplored in the non-stationary domain of…