3 papers
cs.LG2026
A Direct Approach for Handling Contextual Bandits with Latent State Dynamics
Zhen Li, Gilles Stoltz
We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain. We first revisit the simplified model by Nelson et al. (2022),…
stat.ML2025
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
Chiara Mignacco, Matthieu Jonckheere, Gilles Stoltz
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for…
cs.LG2024
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
Daniil Tiapkin, Evgenii Chzhen, Gilles Stoltz
We consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environme…