activity
20192021
most citedRMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents

16 citations · 41 across the 5 of their papers we have counts for

collaborators

6 papers

cs.LG202116 cited

RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents

Wei Qiu, Xinrun Wang, Runsheng Yu +5

Current value-based multi-agent reinforcement learning methods optimize individual Q values to guide individuals' behaviours via centralized training with decentralized execution (…

cs.IR20201 cited

Personalized Adaptive Meta Learning for Cold-start User Preference Prediction

Runsheng Yu, Yu Gong, Xu He +4

A common challenge in personalized user preference prediction is the cold-start problem. Due to the lack of user-item interactions, directly learning from the new users' log data c…

cs.LG20209 cited

Learning to Collaborate in Multi-Module Recommendation via Multi-Agent Reinforcement Learning without Communication

Xu He, Bo An, Yanghua Li +6

With the rise of online e-commerce platforms, more and more customers prefer to shop online. To sell more products, online platforms introduce various modules to recommend items wi…

cs.LG20207 cited

Contextual User Browsing Bandits for Large-Scale Online Mobile Recommendation

Xu He, Bo An, Yanghua Li +4

Online recommendation services recommend multiple commodities to users. Nowadays, a considerable proportion of users visit e-commerce platforms by mobile devices. Due to the limite…

cs.AI20208 cited

Learning Behaviors with Uncertain Human Feedback

Xu He, Haipeng Chen, Bo An

Human feedback is widely used to train agents in many domains. However, previous works rarely consider the uncertainty when humans provide feedback, especially in cases that the op…

cs.AI2019

Learning Efficient Multi-agent Communication: An Information Bottleneck Approach

Rundong Wang, Xu He, Runsheng Yu +3

We consider the problem of the limited-bandwidth communication for multi-agent reinforcement learning, where agents cooperate with the assistance of a communication protocol and a…