5 papers
Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing
Huiling Meng, Ningyuan Chen, Xuefeng Gao
We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space stru…
Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games
Dantong Chu, Xuefeng Gao, Yufei Zhang
As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This pa…
Design Experiments to Compare Multi-armed Bandit Algorithms
Huiling Meng, Ningyuan Chen, Xuefeng Gao
Online platforms routinely compare multi-armed bandit algorithms, such as UCB and Thompson Sampling, to select the best-performing policy. Unlike standard A/B tests for static trea…
Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management
Huiling Meng, Ningyuan Chen, Xuefeng Gao
Intensity control is a class of continuous-time dynamic optimization problems with many important applications in Operations Research including queueing and revenue management. In…
Is Thompson Sampling Susceptible to Algorithmic Collusion?
Yi Xiong, Ningyuan Chen, Xuefeng Gao
When two players are engaged in a repeated game with unknown payoff matrices, they may use single-agent multi-armed bandit algorithms to choose the actions independent of each othe…