40 citations · 125 across the 25 of their papers we have counts for
18 papers · 1 filter
DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning
Canzhe Zhao, Yanjie Ze, Jing Dong +2
Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating…
Player-optimal Stable Regret for Bandit Learning in Matching Markets
Fang Kong, Shuai Li
The problem of matching markets has been studied for a long time in the literature due to its wide range of applications. Finding a stable matching is a common equilibrium objectiv…
Adversarial Attacks on Online Learning to Rank with Click Feedback
Jinhang Zuo, Zhiyao Zhang, Zhiyong Wang +3
Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although p…
Future-conditioned Unsupervised Pretraining for Decision Transformer
Zhihui Xie, Zichuan Lin, Deheng Ye +3
Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…
Best-of-three-worlds Analysis for Linear Bandits with Follow-the-regularized-leader Algorithm
Fang Kong, Canzhe Zhao, Shuai Li
The linear bandit problem has been studied for many years in both stochastic and adversarial settings. Designing an algorithm that can optimize the environment without knowing the…
Efficient Explorative Key-term Selection Strategies for Conversational Contextual Bandits
Zhiyong Wang, Xutong Liu, Shuai Li +1
Conversational contextual bandits elicit user preferences by occasionally querying for explicit feedback on key-terms to accelerate learning. However, there are aspects of existing…