5 papers
Unlearning Offline Stochastic Multi-Armed Bandits
Zichun Ye, Runqi Wang, Xuchuang Wang +3
Machine unlearning aims to unlearn data points from a learned model, offering a principled way to process data-deletion requests and mitigate privacy risks without full retraining.…
Near-Optimal Regret for Efficient Stochastic Combinatorial Semi-Bandits
Zichun Ye, Runqi Wang, Xutong Liu +1
The combinatorial multi-armed bandit (CMAB) is a cornerstone of sequential decision-making framework, dominated by two algorithmic families: UCB-based and adversarial methods such…
Group Distributionally Robust Optimization with Flexible Sample Queries
Haomin Bai, Dingzhi Yu, Shuai Li +2
Group distributionally robust optimization (GDRO) aims to develop models that perform well across distributions simultaneously. Existing GDRO algorithms can only process a fixe…
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
Xutong Liu, Siwei Wang, Jinhang Zuo +7
We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a …
Cascading Bandits Robust to Adversarial Corruptions
Jize Xie, Cheng Chen, Zhiyong Wang +1
Online learning to rank sequentially recommends a small list of items to users from a large candidate set and receives the users' click feedback. In many real-world scenarios, user…