4 papers
Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games
Dantong Chu, Xuefeng Gao, Yufei Zhang
As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This pa…
Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management
Huiling Meng, Ningyuan Chen, Xuefeng Gao
Intensity control is a class of continuous-time dynamic optimization problems with many important applications in Operations Research including queueing and revenue management. In…
Design Experiments to Compare Multi-armed Bandit Algorithms
Huiling Meng, Ningyuan Chen, Xuefeng Gao
Online platforms routinely compare multi-armed bandit algorithms, such as UCB and Thompson Sampling, to select the best-performing policy. Unlike standard A/B tests for static trea…
Is Thompson Sampling Susceptible to Algorithmic Collusion?
Yi Xiong, Ningyuan Chen, Xuefeng Gao
When two players are engaged in a repeated game with unknown payoff matrices, they may use single-agent multi-armed bandit algorithms to choose the actions independent of each othe…