40 citations · 40 across the 4 of their papers we have counts for
4 papers
Learning Adversarial Low-rank Markov Decision Processes with Unknown Transition and Full-information Feedback
Canzhe Zhao, Ruofeng Yang, Baoxiang Wang +2
In this work, we study the low-rank MDPs with adversarially changed losses in the full-information feedback setting. In particular, the unknown transition probability kernel admits…
DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning
Canzhe Zhao, Yanjie Ze, Jing Dong +2
Communication lays the foundation for cooperation in human society and in multi-agent reinforcement learning (MARL). Humans also desire to maintain their privacy when communicating…
Comparison-based Conversational Recommender System with Relative Bandit Feedback
Zhihui Xie, Tong Yu, Canzhe Zhao +1
With the recent advances of conversational recommendations, the recommender system is able to actively and dynamically elicit user preference via conversational interactions. To ac…
Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model
Cheng Chen, Canzhe Zhao, Shuai Li
Online learning to rank (OLTR) interactively learns to choose lists of items from a large collection based on certain click models that describe users' click behaviors. Most recent…