4 papers
Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning
Viktor Veselý, Aleksandar Todorov, Erwan Escudie +1
Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is poorly understood. We identify a…
An -Optimal Sequential Approach for Solving zs-POSGs
Erwan C. Escudie, Matthia Sabatelli, Jilles S. Dibangoye
While recent reductions of zero-sum partially observable stochastic games (zs-POSGs) to transition-independent stochastic games (TI-SGs) theoretically admit dynamic programming, pr…
ε-Optimally Solving Two-Player Zero-Sum POSGs
Erwan Christian Escudie, Matthia Sabatelli, Olivier Buffet +1
We present a novel framework for ε-optimally solving two-player zero-sum partially observable stochastic games (zs-POSGs). These games pose a major challenge due to the absence of…
-Optimally Solving Zero-Sum POSGs
Erwan Escudie, Matthia Sabatelli, Jilles Dibangoye
A recent method for solving zero-sum partially observable stochastic games (zs-POSGs) embeds the original game into a new one called the occupancy Markov game. This reformulation a…