activity
20182022
most citedReinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

25 citations · 25 across the 2 of their papers we have counts for

collaborators

6 papers

stat.ML2022

Phase Transitions in Learning and Earning under Price Protection Guarantee

Qing Feng, Ruihao Zhu, Stefanus Jasin

Motivated by the prevalence of ``price protection guarantee", which allows a customer who purchased a product in the past to receive a refund from the seller during the so-called p…

cs.LG202025 cited

Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions…

cs.LG2019

Non-Stationary Reinforcement Learning: The Blessing of (More) Optimism

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed…

cs.LG2019

Hedging the Drift: Learning to Optimize under Non-Stationarity

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applicatio…

cs.LG2019

Meta Dynamic Pricing: Transfer Learning Across Experiments

Hamsa Bastani, David Simchi-Levi, Ruihao Zhu

We study the problem of learning shared structure \emph{across} a sequence of dynamic pricing experiments for related products. We consider a practical formulation where the unknow…

cs.LG2018

Learning to Route Efficiently with End-to-End Feedback: The Value of Networked Structure

Ruihao Zhu, Eytan Modiano

We introduce efficient algorithms which achieve nearly optimal regrets for the problem of stochastic online shortest path routing with end-to-end feedback. The setting is a natural…