activity
20172024
most citedComparison-based Conversational Recommender System with Relative Bandit Feedback

40 citations · 125 across the 25 of their papers we have counts for

collaborators
Showing cs.LGShow all

18 papers · 1 filter

cs.LG2023

DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning

Canzhe Zhao, Yanjie Ze, Jing Dong +2

Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating…

cs.LG2023

Player-optimal Stable Regret for Bandit Learning in Matching Markets

Fang Kong, Shuai Li

The problem of matching markets has been studied for a long time in the literature due to its wide range of applications. Finding a stable matching is a common equilibrium objectiv…

cs.LG2023

Adversarial Attacks on Online Learning to Rank with Click Feedback

Jinhang Zuo, Zhiyao Zhang, Zhiyong Wang +3

Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although p…

cs.LG2023★ 2 cited

Future-conditioned Unsupervised Pretraining for Decision Transformer

Zhihui Xie, Zichuan Lin, Deheng Ye +3

Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…

cs.LG2023

Best-of-three-worlds Analysis for Linear Bandits with Follow-the-regularized-leader Algorithm

Fang Kong, Canzhe Zhao, Shuai Li

The linear bandit problem has been studied for many years in both stochastic and adversarial settings. Designing an algorithm that can optimize the environment without knowing the…

cs.LG2023

Efficient Explorative Key-term Selection Strategies for Conversational Contextual Bandits

Zhiyong Wang, Xutong Liu, Shuai Li +1

Conversational contextual bandits elicit user preferences by occasionally querying for explicit feedback on key-terms to accelerate learning. However, there are aspects of existing…