3 papers
cs.LG2026
Learning in Markovian bandits with non-observable states and constrained decision epochs
Thomas Hira, Victor Boone, Urtzi Ayesta +1
This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs. The focus is restricte…
cs.IT2025
On the Age of Information in Single-Server Queues with Aged Updates
Fernando Miguelez, Urtzi Ayesta, Josu Doncel +1
The Age of Information (AoI) is a performance metric that quantifies the freshness of data in systems where timely updates are critical. Most state-of-the-art methods typically ass…
math.PR2025
On the instability of local learning algorithms: Q-learning can fail in infinite state spaces
Urtzi Ayesta, Sergey Foss, Matthieu Jonckheere +1
We investigate the challenges of applying model-free reinforcement learning algorithms, like online Q-learning, to infinite state space Markov Decision Processes (MDPs). We first i…