collaborators

11 papers

cs.LG2026

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

Aleksandar Todorov, Matthia Sabatelli

Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value and policy structure may be intr…

cs.LG2026

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

Viktor Veselý, Aleksandar Todorov, Erwan Escudie +1

Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood. We identify a…

cs.GT2026

An -Optimal Sequential Approach for Solving zs-POSGs

Erwan C. Escudie, Matthia Sabatelli, Jilles S. Dibangoye

While recent reductions of zero-sum partially observable stochastic games (zs-POSGs) to transition-independent stochastic games (TI-SGs) theoretically admit dynamic programming, pr…

cs.LG2026

Measuring Orthogonality as the Blind-Spot of Uncertainty Disentanglement

Ivo Pascal de Jong, Andreea Ioana Sburlea, Matthia Sabatelli +1

Aleatoric (data) and epistemic (knowledge) uncertainty are textbook components of Uncertainty Quantification. Jointly estimating these components has been shown to be problematic a…

cs.GT2025

ε-Optimally Solving Two-Player Zero-Sum POSGs

Erwan Christian Escudie, Matthia Sabatelli, Olivier Buffet +1

We present a novel framework for ε-optimally solving two-player zero-sum partially observable stochastic games (zs-POSGs). These games pose a major challenge due to the absence of…

cs.LG2025

On The Presence of Double-Descent in Deep Reinforcement Learning

Viktor Veselý, Aleksandar Todorov, Matthia Sabatelli

The double descent (DD) paradox, where over-parameterized models see generalization improve past the interpolation point, remains largely unexplored in the non-stationary domain of…