5 papers
On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization
Yixuan Zhang, Ruihao Zhu, Qiaomin Xie
Motivated by the principle of satisficing in decision-making, we study satisficing regret guarantees for nonstationary -armed bandits. We show that in the general realizable, pi…
Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards
Wenlong Ji, Yihan Pan, Ruihao Zhu +1
Multi-armed bandit (MAB) is a widely adopted framework for sequential decision-making under uncertainty. Traditional bandit algorithms rely solely on online data, which tends to be…
Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow
Kuo Liang, Yuhang Lu, Jianming Mao +7
Large-scale optimization is a key backbone of modern business decision-making. However, building these models is often labor-intensive and time-consuming. We address this by propos…
Thompson Sampling for Repeated Newsvendor
Li Chen, Hanzhang Qin, Yunbei Xu +2
In this paper, we investigate the performance of Thompson Sampling (TS) for online learning with censored feedback, focusing primarily on the classic repeated newsvendor model--a f…
Contextual Online Pricing with (Biased) Offline Data
Yixuan Zhang, Ruihao Zhu, Qiaomin Xie
We study contextual online pricing with biased offline data. For the scalar price elasticity case, we identify the instance-dependent quantity that measures how far the offl…