2 papers
stat.ML2025
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
Chiara Mignacco, Matthieu Jonckheere, Gilles Stoltz
Online matching problems arise in many complex systems, from cloud services and online marketplaces to organ exchange networks, where timely, principled decisions are critical for…
cs.LG2025
Policy Optimization via Adv2: Adversarial Learning on Advantage Functions
Matthieu Jonckheere, Chiara Mignacco, Gilles Stoltz
We revisit the reduction of learning in adversarial Markov decision processes [MDPs] to adversarial learning based on --values; this reduction has been considered in a number of…