8 papers
Two-Sided Time-Independent Regret for Matching Markets with Limited Interviews
Amirmahdi Mirfakhar, Xuchuang Wang, Mengfan Xu +2
Two-sided matching platforms rely on preferences from both sides, yet participants can evaluate only a small fraction of potential partners. In practice, they use low-cost pre-matc…
Conformal-Style Quantile Analyses for Stochastic Bandits
Chengyu Du, Mengfan Xu
Stochastic bandit algorithms are usually analyzed under a mean-reward criterion, yet many problems favor arms with strong upper-tail performance, which we study herein. For a fixed…
Multi-Objective Multi-Agent Bandits: From Learning Efficiency to Fairness Optimization
John Wang, Mengfan Xu
We study multi-objective multi-agent multi-armed bandits (MO-MA-MAB) under stochastic rewards, where agents observe heterogeneous reward vectors and communicate over time-varying g…
Bandit Learning in General Open Multi-agent Systems
Mengfan Xu
Recent developments in digital platforms have highlighted the prevalence of open systems, where agents can arrive and depart over time. While bandit learning in open systems has re…
Are Stochastic Multi-objective Bandits Harder than Single-objective Bandits?
Changkun Guan, Mengfan Xu
Multi-objective bandits have attracted increasing attention for their broad applicability, with \(d\)-dimensional reward vectors inducing Pareto regret. There has been a subtle deb…
Distributed Multi-Agent Bandits Over ErdÅs-Rényi Random Networks
Jingyuan Liu, Hao Qiu, Lin Yang +1
We study the distributed multi-agent multi-armed bandit problem with heterogeneous rewards over random communication graphs. Uniquely, at each time step agents communicate over…