activity
20192026
most citedReinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

25 citations · 40 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG202025 cited

Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions…

cs.LG2020

Best Arm Identification for Cascading Bandits in the Fixed Confidence Setting

Zixin Zhong, Wang Chi Cheung, Vincent Y. F. Tan

We design and analyze CascadeBAI, an algorithm for finding the best set of items, also called an arm, within the framework of cascading bandits. An upper bound on the time comp…

cs.LG2019

Non-Stationary Reinforcement Learning: The Blessing of (More) Optimism

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed…

cs.LG201911 cited

Exploration-Exploitation Trade-off in Reinforcement Learning on Online Markov Decision Processes with Global Concave Rewards

Wang Chi Cheung

We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on th…

cs.LG2019

Hedging the Drift: Learning to Optimize under Non-Stationarity

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applicatio…