Publications (84)
Online Learning and Optimization for Revenue Management Problems with Add-on Discounts
David Simchi-Levi, Rui Sun, Huanan Zhang
We study in this paper a revenue management problem with add-on discounts. The problem is motivated by the practice in the video game industry, where a retailer offers discounts on…
PPI-SVRG: Unifying Prediction-Powered Inference and Variance Reduction for Semi-Supervised Optimization
Ruicheng Ao, Hongyu Chen, Haoyang Liu +2
We study semi-supervised stochastic optimization when labeled data is scarce but predictions from pre-trained models are available. PPI and SVRG both reduce variance through contro…
Non-Stationary Reinforcement Learning: The Blessing of (More) Optimism
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under temporal drifts, ie, both the reward and state transition distributions are allowed…
Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
Dylan J. Foster, Akshay Krishnamurthy, David Simchi-Levi +1
We consider the offline reinforcement learning problem, where the aim is to learn a decision making policy from logged data. Offline RL -- particularly when coupled with (value) fu…
The Lingering of Gradients: Theory and Applications
Zeyuan Allen-Zhu, David Simchi-Levi, Xinshang Wang
Classically, the time complexity of a first-order method is estimated by its number of gradient computations. In this paper, we study a more refined complexity by taking into accou…
Hedging the Drift: Learning to Optimize under Non-Stationarity
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We introduce data-driven decision-making algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary bandit settings. These settings capture applicatio…
Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations
Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao +2
Negotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reser…
Perturbing the Derivative: Doubly Wild Refitting for Model-Free Evaluation of Opaque Machine Learning Predictors
Haichen Hu, David Simchi-Levi
We study the problem of excess risk evaluation for empirical risk minimization (ERM) under convex losses. We show that by leveraging the idea of wild refitting, one can upper bound…
Optimal Adaptive Experimental Design for Estimating Treatment Effect
Jiachun Li, David Simchi-Levi, Yunxiao Zhao
Given n experiment subjects with potentially heterogeneous covariates and two possible treatments, namely active treatment and control, this paper addresses the fundamental questio…
Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles
Haichen Hu, David Simchi-Levi
We study whether stochastic nonconvex optimization can be reduced to ordinary static regret minimization in online convex optimization in a black-box manner. For smooth nonconvex o…
Sobolev Norm Learning Rates for Conditional Mean Embeddings
Prem Talwai, Ali Shameli, David Simchi-Levi
We develop novel learning rates for conditional mean embeddings by applying the theory of interpolation for reproducing kernel Hilbert spaces (RKHS). We derive explicit, adaptive c…
LLM Evaluation as Tensor Completion: Low Rank Structure and Semiparametric Efficiency
Jiachun Li, David Simchi-Levi, Will Wei Sun
Large language model (LLM) evaluation platforms increasingly rely on pairwise human judgments. These data are noisy, sparse, and non-uniform, yet leaderboards are reported with lim…
Refined Thompson Learning for Adaptive Bandits: Sustainable Power-Efficient Flexibility Scheduling Across Data Centers
Yifu Ding, Zixi Chen, Ruicheng Ao +2
The rapid rise in energy consumption from large-scale AI workloads in data centers placed the increasing pressures on power grids in recent years. Since grids must maintain real-ti…
Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
Jian Qian, Haichen Hu, David Simchi-Levi
Motivated by the recent discovery of a statistical and computational reduction from contextual bandits to offline regression (Simchi-Levi and Xu, 2021), we address the general (sto…
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research
Ruicheng Ao, David Simchi-Levi, Xinshang Wang
Operations Research practitioners debug infeasible models through an iterative process: inspecting Irreducible Infeasible Subsystems ( IIS), identifying constraint conflicts, and r…
Resource-Constrained Adaptive Inference for Sequential Pricing
Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
Resource-constrained pricing controllers can make fixed-price inference impossible: the controller's resource state may remove the target price neighborhood from the feasible set,…
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
Rui Ai, Yu Pan, David Simchi-Levi +1
In user-agent interaction scenarios such as recommendation, brainstorming, and code suggestion, Large Language Models (LLMs) often generate sets of candidate recommendations where…
Meta Dynamic Pricing: Transfer Learning Across Experiments
Hamsa Bastani, David Simchi-Levi, Ruihao Zhu
We study the problem of learning shared structure \emph{across} a sequence of dynamic pricing experiments for related products. We consider a practical formulation where the unknow…
On Policies for Single-leg Revenue Management with Limited Demand Information
Will Ma, David Simchi-Levi, Chung-Piaw Teo
In this paper we study the single-item revenue management problem, with no information given about the demand trajectory over time. When the item is sold through accepting/rejectin…
On the Reliability Limits of LLM-Based Multi-Agent Planning
Ruicheng Ao, Siyang Gao, David Simchi-Levi
This technical note studies the reliability limits of LLM-based multi-agent planning as a delegated decision problem. We model the LLM-based multi-agent architecture as a finite ac…
Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons
Jiachun Li, David Simchi-Levi, Will Wei Sun
Pairwise human-preference platforms such as Chatbot Arena have become central to large language model (LLM) evaluation, yet reliable task-specific ranking remains challenging. Glob…
Online Pricing with Offline Data: Phase Transition and Inverse Square Law
Jinzhi Bu, David Simchi-Levi, Yunzong Xu
This paper investigates the impact of pre-existing offline data on online learning, in the context of dynamic pricing. We study a single-product dynamic pricing problem over a sell…
Constrained Online Decision-Making: A Unified Framework
Haichen Hu, David Simchi-Levi, Navid Azizan
Contextual online decision-making problems with constraints appear in a wide range of real-world applications, such as adaptive experimental design under safety constraints, person…
Perturbing the Derivative: Wild Refitting for Model-Free Evaluation of Machine Learning Models under Bregman Losses
Haichen Hu, David Simchi-Levi
We study the excess risk evaluation of classical penalized empirical risk minimization (ERM) with Bregman losses. We show that by leveraging the idea of wild refitting, one can eff…
Prediction-Guided Active Experiments
Ruicheng Ao, Hongyu Chen, David Simchi-Levi
In this work, we introduce a new framework for active experimentation, the Prediction-Guided Active Experiment (PGAE), which leverages predictions from an existing machine learning…
The Value of Information in Resource-Constrained Pricing
Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
Firms that price perishable resources -- airline seats, hotel rooms, seasonal inventory -- now routinely use demand predictions, but these predictions vary widely in quality. Under…
Inventory Balancing with Online Learning
Wang Chi Cheung, Will Ma, David Simchi-Levi +1
We study a general problem of allocating limited resources to heterogeneous customers over time under model uncertainty. Each type of customer can be serviced using different actio…
Phase Transitions in Bandits with Switching Constraints
David Simchi-Levi, Yunzong Xu
We consider the classical stochastic multi-armed bandit problem with a constraint that limits the total cost incurred by switching between actions to be no larger than a given swit…
Contextual Online Decision Making with Infinite-Dimensional Functional Regression
Haichen Hu, Rui Ai, Stephen Bates +1
Contextual sequential decision-making problems play a crucial role in machine learning, encompassing a wide range of downstream applications such as bandits, sequential hypothesis…
Utility Fairness in Contextual Dynamic Pricing with Demand Learning
Xi Chen, David Simchi-Levi, Yining Wang
This paper introduces a novel contextual bandit algorithm for personalized pricing under utility fairness constraints in scenarios with uncertain demand, achieving an optimal regre…
Shrinking the Upper Confidence Bound: A Dynamic Product Selection Problem for Urban Warehouses
Rong Jin, David Simchi-Levi, Li Wang +2
The recent rising popularity of ultra-fast delivery services on retail platforms fuels the increasing use of urban warehouses, whose proximity to customers makes fast deliveries vi…
Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
Rui Ai, Yuqi Pan, David Simchi-Levi +2
With the rapid progress of multi-agent large language model (LLM) reasoning, how to effectively aggregate answers from multiple LLMs has emerged as a fundamental challenge. Standar…
Learning to Optimize under Non-Stationarity
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We introduce algorithms that achieve state-of-the-art \emph{dynamic regret} bounds for non-stationary linear stochastic bandit setting. It captures natural applications such as dyn…
Multi-agent Adaptive Mechanism Design
Qiushi Han, David Simchi-Levi, Renfei Tan +1
We study a sequential mechanism design problem in which a principal seeks to elicit truthful reports from multiple rational agents while starting with no prior knowledge of agents'…
Privacy Preserving Adaptive Experiment Design
Jiachun Li, Kaining Shi, David Simchi-Levi
Adaptive experiment is widely adopted to estimate conditional average treatment effect (CATE) in clinical trials and many other scenarios. While the primary goal in experiment is t…
Solve Smart, Not Often: Policy Learning for Costly MILP Re-solving
Rui Ai, Hugo De Oliveira Barbalho, Sirui Li +3
A common challenge in real-time operations is deciding whether to re-solve an optimization problem or continue using an existing solution. While modern data platforms may collect i…
Dynamic Pricing and Demand Learning on a Large Network of Products: A PAC-Bayesian Approach
N. Bora Keskin, David Simchi-Levi, Prem Talwai
We consider a seller offering a large network of products over a time horizon of periods. The seller does not know the parameters of the products' linear demand model, and…
Regret Distribution in Stochastic Bandits: Optimal Trade-off between Expectation and Tail Risk
David Simchi-Levi, Zeyu Zheng, Feng Zhu
We study the optimal trade-off between expectation and tail risk for regret distribution in the stochastic multi-armed bandit model. We fully characterize the interplay among three…
Designing Service Systems from Textual Evidence
Ruicheng Ao, Hongyu Chen, Siyang Gao +2
Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing policy, or the most effective quality contro…
Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control
Weichao Mao, Kaiqing Zhang, Ruihao Zhu +2
We consider model-free reinforcement learning (RL) in non-stationary Markov decision processes. Both the reward functions and the state transition functions are allowed to vary arb…
Design and Analysis of Switchback Experiments
Iavor Bojinov, David Simchi-Levi, Jinglong Zhao
Switchback experiments, where a firm sequentially exposes an experimental unit to random treatments, are among the most prevalent designs used in the technology sector, with applic…
Learning to Price with Resource Constraints: From Full Information to Machine-Learned Prices
Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
We study the dynamic pricing problem with knapsack, addressing the challenge of balancing exploration and exploitation under resource constraints. We introduce three algorithms tai…
Semiparametric Efficiency in Sequential Experiments: Characterization and Design via Average Propensity
Jiachun Li, David Simchi-Levi
Modern experiments, including evaluations of AI-enabled services and platform interventions, often depart from independent and identically distributed (i.i.d.) sampling because ass…
Reaping the Benefits of Bundling under High Production Costs
Will Ma, David Simchi-Levi
It is well-known that selling different goods in a single bundle can significantly increase revenue. However, bundling is no longer profitable if the goods have high production cos…
Nonparametric Regression in Dirichlet Spaces: A Random Obstacle Approach
Prem Talwai, David Simchi-Levi
In this paper, we consider nonparametric estimation over general Dirichlet metric measure spaces. Unlike the more commonly studied reproducing kernel Hilbert space, whose elements…
Assortment Optimization under Unknown MultiNomial Logit Choice Models
Wang Chi Cheung, David Simchi-Levi
Motivated by e-commerce, we study the online assortment optimization problem. The seller offers an assortment, i.e. a subset of products, to each arriving customer, who then purcha…
Privacy-Preserving Dynamic Personalized Pricing with Demand Learning
Xi Chen, David Simchi-Levi, Yining Wang
The prevalence of e-commerce has made detailed customers' personal information readily accessible to retailers, and this information has been widely used in pricing decisions. When…
Blind Network Revenue Management and Bandits with Knapsacks under Limited Switches
David Simchi-Levi, Yunzong Xu, Jinglong Zhao
This paper studies the impact of limited switches on resource-constrained dynamic pricing with demand learning. We focus on the classical price-based blind network revenue manageme…
From Confounding to Learning: Dynamic Service Fee Pricing on Third-Party Platforms
Rui Ai, David Simchi-Levi, Feng Zhu
We study the pricing behavior of third-party platforms facing strategic agents. Assuming the platform is a revenue maximizer, it observes market features that generally affect dema…
On the Optimal Regret of Locally Private Linear Contextual Bandit
Jiachun Li, David Simchi-Levi, Yining Wang
Contextual bandit with linear reward functions is among one of the most extensively studied models in bandit and online learning research. Recently, there has been increasing inter…
Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation
Haichen Hu, Jian Qian, David Simchi-Levi
Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms require repeated, costly calls…
Pre-Trained AI Model Assisted Online Decision-Making under Missing Covariates: A Theoretical Perspective
Haichen Hu, David Simchi-Levi
We study a sequential contextual decision-making problem in which certain covariates are missing but can be imputed using a pre-trained AI model. From a theoretical perspective, we…
Large Language Models for Supply Chain Decisions
David Simchi-Levi, Konstantina Mellou, Ishai Menache +1
Supply Chain Management requires addressing a variety of complex decision-making challenges, from sourcing strategies to planning and execution. Over the last few decades, advances…
A Simple and Optimal Policy Design with Safety against Heavy-Tailed Risk for Stochastic Bandits
David Simchi-Levi, Zeyu Zheng, Feng Zhu
We study the stochastic multi-armed bandit problem and design new policies that enjoy both worst-case optimality for expected regret and light-tailed risk for regret distribution.…
Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism
Wang Chi Cheung, David Simchi-Levi, Ruihao Zhu
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions…
Online Resource Allocation with Average Budget Constraints
Ruicheng Ao, Hongyu Chen, David Simchi-Levi +1
We consider the problem of online resource allocation with average budget constraints. At each time point the decision maker makes an irrevocable decision of whether to accept or r…
Bypassing the Monster: A Faster and Simpler Optimal Algorithm for Contextual Bandits under Realizability
David Simchi-Levi, Yunzong Xu
We consider the general (stochastic) contextual bandit problem under the realizability assumption, i.e., the expected reward, as a function of contexts and actions, belongs to a ge…
Interleaved Resampling and Refitting: Data and Compute-Efficient Evaluation of Black-Box Predictors
Haichen Hu, David Simchi-Levi
We study the problem of evaluating the excess risk of large-scale empirical risk minimization under the square loss. Leveraging the idea of wild refitting and resampling, we assume…
Algorithms for Online Matching, Assortment, and Pricing with Tight Weight-dependent Competitive Ratios
Will Ma, David Simchi-Levi
Motivated by the dynamic assortment offerings and item pricings occurring in e-commerce, we study a general problem of allocating finite inventories to heterogeneous customers arri…
Bayesian Mechanism Design for Blockchain Transaction Fee Allocation
Xi Chen, David Simchi-Levi, Zishuo Zhao +1
In blockchain systems, the design of transaction fee mechanisms is essential for stability and satisfaction for both miners and users. A recent work has proven the impossibility of…
Designing Information Delays in Supply Chains
Prem Talwai, Rene Caldentey, Avi Giloni +3
This paper studies how a downstream retailer in a decentralized two-tier supply chain can implicitly transmit demand information to an upstream supplier through the structure of it…
Offline Planning and Online Learning under Recovering Rewards
David Simchi-Levi, Zeyu Zheng, Feng Zhu
Motivated by emerging applications such as live-streaming e-commerce, promotions and recommendations, we introduce and solve a general class of non-stationary multi-armed bandit pr…
Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
Dylan J. Foster, Alexander Rakhlin, David Simchi-Levi +1
In the classical multi-armed bandit problem, instance-dependent algorithms attain improved performance on "easy" problems with a gap between the best and second-best arm. Are simil…
OptiRepair: Closed-Loop Diagnosis and Repair of Supply Chain Optimization Models with LLM Agents
Ruicheng Ao, David Simchi-Levi, Xinshang Wang
Supply chain optimization models frequently become infeasible because of modeling errors. Diagnosis and repair require scarce OR expertise: analysts must interpret solver diagnosti…
Beyond ATE: Multi-Criteria Design for A/B Testing
Jiachun Li, Kaining Shi, David Simchi-Levi
In the era of large-scale AI deployment and high-stakes clinical trials, adaptive experimentation faces a ``trilemma'' of conflicting objectives: minimizing cumulative regret (welf…
A Practically Competitive and Provably Consistent Algorithm for Uplift Modeling
Yan Zhao, Xiao Fang, David Simchi-Levi
Randomized experiments have been critical tools of decision making for decades. However, subjects can show significant heterogeneity in response to treatments in many important app…
GenAI vs. Human Creators: Procurement Mechanism Design in Two-/Three-Layer Markets
Rui Ai, David Simchi-Levi, Haifeng Xu
With the rapid advancement of generative AI (GenAI), mechanism design adapted to its unique characteristics poses new theoretical and practical challenges. Unlike traditional goods…
Multi-stage and Multi-customer Assortment Optimization with Inventory Constraints
Elaheh Fata, Will Ma, David Simchi-Levi
We consider an assortment optimization problem where a customer chooses a single item from a sequence of sets shown to her, while limited inventories constrain the items offered to…
Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
Jiachun Li, David Simchi-Levi
Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficiency. The oracle design is a covariate-depe…
Dynamic Pricing (and Assortment) under a Static Calendar
Will Ma, David Simchi-Levi, Jinglong Zhao
This work is motivated by our collaboration with a large consumer packaged goods (CPG) company. We have found that while the company appreciates the advantages of dynamic pricing,…
Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards
Shuze Daniel Liu, Claire Chen, Jiabao Sean Xiao +4
The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomple…
Provably More Efficient Q-Learning in the One-Sided-Feedback/Full-Feedback Settings
Xiao-Yue Gong, David Simchi-Levi
Motivated by the episodic version of the classical inventory control problem, we propose a new Q-learning-based algorithm, Elimination-Based Half-Q-Learning (HQL), that enjoys impr…
Improving the Estimation of Lifetime Effects in A/B Testing via Treatment Locality
Shuze Chen, David Simchi-Levi, Chonghuan Wang
Utilizing randomized experiments to evaluate the effect of short-term treatments on the short-term outcomes has been well understood and become the golden standard in industrial pr…
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
Ruicheng Ao, Gan Luo, David Simchi-Levi +1
Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires token-by-token inference, making GPU sched…
Uplift Modeling with Multiple Treatments and General Response Types
Yan Zhao, Xiao Fang, David Simchi-Levi
Randomized experiments have been used to assist decision-making in many areas. They help people select the optimal treatment for the test population with certain statistical guaran…
When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play
Junyi Sha, Renfei Tan, David Simchi-Levi
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks, but its effect on behavioral diversity in sequential decision-making remains under-e…
Assured autonomy: How operations research powers and orchestrates generative AI systems
Tinglong Dai, David Simchi-Levi, Michelle Xiao Wu +1
Generative artificial intelligence (GenAI) is shifting from conversational assistants toward agentic systems -- autonomous decision-making systems that sense, decide, and act withi…
Partial Identification under Missing Data Using Weak Shadow Variables from Pretrained Models
Hongyu Chen, David Simchi-Levi, Ruoxuan Xiong
Estimating population quantities such as mean outcomes from user feedback is fundamental to platform evaluation and social science, yet feedback is often missing not at random (MNA…
Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
Carol Xuan Long, David Simchi-Levi, Feng Zhu +3
This paper studies autonomous generative AI agents in multi-echelon supply chains using the MIT Beer Game. We identify four inference-time levers that shape performance: model sele…
The Competitive Ratio of Threshold Policies for Online Unit-density Knapsack Problems
Will Ma, David Simchi-Levi, Jinglong Zhao
We study a wholesale supply chain ordering problem. In this problem, the supplier has an initial stock, and faces an unpredictable stream of incoming orders, making real-time decis…
Beyond Covariance Matrix: The Statistical Complexity of Private Linear Regression
Fan Chen, Jiachun Li, Alexander Rakhlin +1
We study the statistical complexity of private linear regression under an unknown, potentially ill-conditioned covariate distribution. Somewhat surprisingly, under privacy constrai…
Optimal Learning Rates for Regularized Least-Squares with a Fourier Capacity Condition
Prem Talwai, David Simchi-Levi
We derive minimax adaptive rates for a new, broad class of Tikhonov-regularized learning problems in Hilbert scales under general source conditions. Our analysis does not require t…
Best Arm Identification with LLM Judges and Limited Human
Ruicheng Ao, Hongyu Chen, Siyang Gao +2
We study fixed-confidence best-arm identification (BAI) where a cheap but potentially biased proxy (e.g., LLM judge) is available for every sample, while an expensive ground-truth…
Service-Induced Congestion in Memory-Constrained LLM Serving
Ruicheng Ao, Jing Dong, Gan Luo +1
In large language model (LLM) serving, each request accumulates persistent graphics processing unit (GPU) memory during service as its key-value cache grows with every generated to…