collaborators

5 papers

cs.LG2026

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

Huiling Meng, Ningyuan Chen, Xuefeng Gao

We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space stru…

math.OC2026

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

Dantong Chu, Xuefeng Gao, Yufei Zhang

As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This pa…

cs.LG2026

Design Experiments to Compare Multi-armed Bandit Algorithms

Huiling Meng, Ningyuan Chen, Xuefeng Gao

Online platforms routinely compare multi-armed bandit algorithms, such as UCB and Thompson Sampling, to select the best-performing policy. Unlike standard A/B tests for static trea…

cs.LG2024

Reinforcement Learning for Intensity Control: An Application to Choice-Based Network Revenue Management

Huiling Meng, Ningyuan Chen, Xuefeng Gao

Intensity control is a class of continuous-time dynamic optimization problems with many important applications in Operations Research including queueing and revenue management. In…

cs.GT2024

Is Thompson Sampling Susceptible to Algorithmic Collusion?

Yi Xiong, Ningyuan Chen, Xuefeng Gao

When two players are engaged in a repeated game with unknown payoff matrices, they may use single-agent multi-armed bandit algorithms to choose the actions independent of each othe…