4 papers
A Direct Approach for Handling Contextual Bandits with Latent State Dynamics
Zhen Li, Gilles Stoltz
We consider a linear contextual bandit model where contexts and rewards are governed by a finite hidden Markov chain. We first revisit the simplified model by Nelson et al. (2022),…
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
Chiara Mignacco, Matthieu Jonckheere, Gilles Stoltz
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for…
Policy Optimization via Adv2: Adversarial Learning on Advantage Functions
Matthieu Jonckheere, Chiara Mignacco, Gilles Stoltz
We revisit the reduction of learning in adversarial Markov decision processes [MDPs] to adversarial learning based on --values; this reduction has been considered in a number of…
Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization
Daniil Tiapkin, Evgenii Chzhen, Gilles Stoltz
We consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environme…