activity
20222026
collaborators

6 papers

stat.ML2026

On the Peril of (Even a Little) Nonstationarity in Satisficing Regret Minimization

Yixuan Zhang, Ruihao Zhu, Qiaomin Xie

Motivated by the principle of satisficing in decision-making, we study satisficing regret guarantees for nonstationary -armed bandits. We show that in the general realizable, pi…

cs.AI2026

Large-Scale Optimization Model Auto-Formulation: Harnessing LLM Flexibility via Structured Workflow

Kuo Liang, Yuhang Lu, Jianming Mao +7

Large-scale optimization is a key backbone of modern business decision-making. However, building these models is often labor-intensive and time-consuming. We address this by propos…

cs.LG2025

Contextual Online Pricing with (Biased) Offline Data

Yixuan Zhang, Ruihao Zhu, Qiaomin Xie

We study contextual online pricing with biased offline data. For the scalar price elasticity case, we identify the instance-dependent quantity that measures how far the offli…

math.ST2025

Multi-Armed Bandits With Machine Learning-Generated Surrogate Rewards

Wenlong Ji, Yihan Pan, Ruihao Zhu +1

Multi-armed bandit (MAB) is a widely adopted framework for sequential decision-making under uncertainty. Traditional bandit algorithms rely solely on online data, which tends to be…

cs.LG2025

Thompson Sampling for Repeated Newsvendor

Li Chen, Hanzhang Qin, Yunbei Xu +2

In this paper, we investigate the performance of Thompson Sampling (TS) for online learning with censored feedback, focusing primarily on the classic repeated newsvendor model--a f…

cs.LG2022

Learning to Price Supply Chain Contracts against a Learning Retailer

Xuejun Zhao, Ruihao Zhu, William B. Haskell

The rise of big data analytics has automated the decision-making of companies and increased supply chain agility. In this paper, we study the supply chain contract design problem f…