25 citations · 33 across the 8 of their papers we have counts for
8 papers · 1 filter
Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi +1
In the classical multi-armed bandit problem, instance-dependent algorithms attain improved performance on "easy" problems with a gap between the best and second-best arm. Are simil…
Provably More Efficient Q-Learning in the One-Sided-Feedback/Full-Feedback Settings
Xiao-Yue Gong, David Simchi-Levi
Motivated by the episodic version of the classical inventory control problem, we propose a new Q-learning-based algorithm, Elimination-Based Half-Q-Learning (HQL), that enjoys impr…
Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions…
Non-Stationary Reinforcement Learning: The Blessing of (More) Optimism
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed…
Phase Transitions in Bandits with Switching Constraints
David Simchi-Levi, Yunzong Xu
We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given swit…
Hedging the Drift: Learning to Optimize under Non-Stationarity
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applicatio…