5 papers
Optimism Stabilizes Thompson Sampling for Adaptive Inference
Shunxing Yan, Han Zhong
Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are subtle. Classical asymptotic theory for…
Robust Assortment Optimization from Observational Data
Miao Lu, Yuxuan Han, Han Zhong +2
Assortment optimization is a fundamental challenge in modern retail and recommendation systems, where the goal is to select a subset of products that maximizes expected revenue und…
Learning an Optimal Assortment Policy under Observational Data
Yuxuan Han, Han Zhong, Miao Lu +2
We study the fundamental problem of offline assortment optimization under the Multinomial Logit (MNL) model, where sellers must determine the optimal subset of the products to offe…
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
Xutong Liu, Siwei Wang, Jinhang Zuo +7
We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a …
A3S: A General Active Clustering Method with Pairwise Constraints
Xun Deng, Junlong Liu, Han Zhong +5
Active clustering aims to boost the clustering performance by integrating human-annotated pairwise constraints through strategic querying. Conventional approaches with semi-supervi…