2 citations · 2 across the 3 of their papers we have counts for
4 papers
Cascading Bandits Robust to Adversarial Corruptions
Jize Xie, Cheng Chen, Zhiyong Wang +1
Online learning to rank sequentially recommends a small list of items to users from a large candidate set and receives the users' click feedback. In many real-world scenarios, user…
Combinatorial Multivariant Multi-Armed Bandits with Applications to Episodic Reinforcement Learning and Beyond
Xutong Liu, Siwei Wang, Jinhang Zuo +7
We introduce a novel framework of combinatorial multi-armed bandits (CMAB) with multivariant and probabilistically triggering arms (CMAB-MT), where the outcome of each arm is a …
Online Corrupted User Detection and Regret Minimization
Zhiyong Wang, Jize Xie, Tong Yu +2
In real-world online web systems, multiple users usually arrive sequentially into the system. For applications like click fraud and fake reviews, some users can maliciously perform…
Online Clustering of Bandits with Misspecified User Models
Zhiyong Wang, Jize Xie, Xutong Liu +2
The contextual linear bandit is an important online learning problem where given arm features, a learning agent selects an arm at each round to maximize the cumulative rewards in t…