activity
20172023
most citedDynamic Regret of Policy Optimization in Non-stationary Environments

11 citations · 28 across the 7 of their papers we have counts for

collaborators

11 papers

cs.PF20201 cited

Zero Queueing for Multi-Server Jobs

Weina Wang, Qiaomin Xie, Mor Harchol-Balter

Cloud computing today is dominated by multi-server jobs. These are jobs that request multiple servers simultaneously and hold onto all of these servers for the duration of the job.…

cs.LG2020

Provable Fictitious Play for General Mean-Field Games

Qiaomin Xie, Zhuoran Yang, Zhaoran Wang +1

We propose a reinforcement learning algorithm for stationary mean-field games, where the goal is to learn a pair of mean-field state and stationary policy that constitutes the Nash…

cs.LG202011 cited

Dynamic Regret of Policy Optimization in Non-stationary Environments

Yingjie Fei, Zhuoran Yang, Zhaoran Wang +1

We consider reinforcement learning (RL) in episodic MDPs with adversarial full-information reward feedback and unknown fixed transition kernels. We propose two model-free policy op…

cs.LG20209 cited

Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret

Yingjie Fei, Zhuoran Yang, Yudong Chen +2

We study risk-sensitive reinforcement learning in episodic Markov decision processes with unknown transition kernels, where the goal is to optimize the total reward under the risk…

cs.LG20203 cited

Stable Reinforcement Learning with Unbounded State Space

Devavrat Shah, Qiaomin Xie, Zhi Xu

We consider the problem of reinforcement learning (RL) with unbounded state space motivated by the classical problem of scheduling in a queueing network. Traditional policies as we…

cs.AI2020

POLY-HOOT: Monte-Carlo Planning in Continuous Space MDPs with Non-Asymptotic Analysis

Weichao Mao, Kaiqing Zhang, Qiaomin Xie +1

Monte-Carlo planning, as exemplified by Monte-Carlo Tree Search (MCTS), has demonstrated remarkable performance in applications with finite spaces. In this paper, we consider Monte…