2 citations · 2 across the 3 of their papers we have counts for
3 papers
stat.ML2023
Contextual Bandits for Evaluating and Improving Inventory Control Policies
Dean Foster, Randy Jia, Dhruv Madeka
Solutions to address the periodic review inventory control problem with nonstationary random demand, lost sales, and stochastic vendor lead times typically involve making strong as…
cs.LG2023
Learning an Inventory Control Policy with General Inventory Arrival Dynamics
Sohrab Andaz, Carson Eisenach, Dhruv Madeka +4
In this paper we address the problem of learning and backtesting inventory control policies in the presence of general arrival dynamics -- which we term as a quantity-over-time arr…
cs.LG2017★ 2 cited
Posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal, Randy Jia
We present an algorithm based on posterior sampling (aka Thompson sampling) that achieves near-optimal worst-case regret bounds when the underlying Markov Decision Process (MDP) is…