activity
20202023
most citedDecentralized Multi-player Multi-armed Bandits with No Collision Information

12 citations · 24 across the 6 of their papers we have counts for

collaborators

7 papers

stat.ML2023

Provably Efficient Offline Reinforcement Learning with Perturbed Data Sources

Chengshuai Shi, Wei Xiong, Cong Shen +1

Existing theoretical studies on offline reinforcement learning (RL) mostly consider a dataset sampled directly from the target task. In practice, however, data often come from seve…

cs.LG2022

A Self-Play Posterior Sampling Algorithm for Zero-Sum Markov Games

Wei Xiong, Han Zhong, Chengshuai Shi +2

Existing studies on provably efficient algorithms for Markov games (MGs) almost exclusively build on the "optimism in the face of uncertainty" (OFU) principle. This work focuses on…

stat.ML20216 cited

Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and Generalization

Chengshuai Shi, Wei Xiong, Cong Shen +1

Despite the significant interests and many progresses in decentralized multi-player multi-armed bandits (MP-MAB) problems in recent years, the regret gap to the natural centralized…

stat.ML20212 cited

(Almost) Free Incentivized Exploration from Decentralized Learning Agents

Chengshuai Shi, Haifeng Xu, Wei Xiong +1

Incentivized exploration in multi-armed bandits (MAB) has witnessed increasing interests and many progresses in recent years, where a principal offers bonuses to agents to do explo…

cs.LG20214 cited

Distributional Reinforcement Learning for Multi-Dimensional Reward Functions

Pushi Zhang, Xiaoyu Chen, Li Zhao +3

A growing trend for value-based reinforcement learning (RL) algorithms is to capture more information than scalar value functions in the value network. One of the most well-known m…

math.OC2020

PMGT-VR: A decentralized proximal-gradient algorithmic framework with variance reduction

Haishan Ye, Wei Xiong, Tong Zhang

This paper considers the decentralized composite optimization problem. We propose a novel decentralized variance-reduction proximal-gradient algorithmic framework, called PMGT-VR,…