activity
20172020
most citedReinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

25 citations · 33 across the 8 of their papers we have counts for

collaborators
Showing 2019Show all

8 papers · 1 filter

cs.DS2019

Multi-stage and Multi-customer Assortment Optimization with Inventory Constraints

Elaheh Fata, Will Ma, David Simchi-Levi

We consider an assortment optimization problem where a customer chooses a single item from a sequence of sets shown to her, while limited inventories constrain the items offered to…

cs.LG2019

Non-Stationary Reinforcement Learning: The Blessing of (More) Optimism

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed…

cs.DS20191 cited

Algorithms for Online Matching, Assortment, and Pricing with Tight Weight-dependent Competitive Ratios

Will Ma, David Simchi-Levi

Motivated by the dynamic assortment offerings and item pricings occurring in e-commerce, we study a general problem of allocating finite inventories to heterogeneous customers arri…

math.OC20191 cited

Shrinking the Upper Confidence Bound: A Dynamic Product Selection Problem for Urban Warehouses

Rong Jin, David Simchi-Levi, Li Wang +2

The recent rising popularity of ultra-fast delivery services on retail platforms fuels the increasing use of urban warehouses, whose proximity to customers makes fast deliveries vi…

cs.LG2019

Phase Transitions in Bandits with Switching Constraints

David Simchi-Levi, Yunzong Xu

We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given swit…

cs.LG2019

Hedging the Drift: Learning to Optimize under Non-Stationarity

Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu

We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applicatio…