activity
20212024
most citedLearning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

4 citations · 16 across the 8 of their papers we have counts for

collaborators

8 papers

cs.AI20241 cited

Multi-Agent Reinforcement Learning with a Hierarchy of Reward Machines

Xuejing Zheng, Chao Yu

In this paper, we study the cooperative Multi-Agent Reinforcement Learning (MARL) problems using Reward Machines (RMs) to specify the reward functions such that the prior knowledge…

cs.MA2023

Models as Agents: Optimizing Multi-Step Predictions of Interactive Local Models in Model-Based Multi-Agent Reinforcement Learning

Zifan Wu, Chao Yu, Chen Chen +2

Research in model-based reinforcement learning has made significant progress in recent years. Compared to single-agent settings, the exponential dimension growth of the joint state…

eess.SP20231 cited

mmAlert: mmWave Link Blockage Prediction via Passive Sensing

Chao Yu, Yifei Sun, Yan Luo +1

In this letter, the mmAlert system, predicting millimeter wave (mmWave) link blockage during data communication, is elaborated and demonstrated. The passive sensing method is adopt…

cs.RO20233 cited

Learning Graph-Enhanced Commander-Executor for Multi-Agent Navigation

Xinyi Yang, Shiyu Huang, Yiwen Sun +5

This paper investigates the multi-agent navigation problem, which requires multiple agents to reach the target goals in a limited time. Multi-agent reinforcement learning (MARL) ha…

cs.AI20234 cited

Learning Zero-Shot Cooperation with Humans, Assuming Humans Are Biased

Chao Yu, Jiaxuan Gao, Weilin Liu +5

There is a recent trend of applying multi-agent reinforcement learning (MARL) to train an agent that can cooperate with humans in a zero-shot fashion without using any human data.…

cs.LG20234 cited

Plan To Predict: Learning an Uncertainty-Foreseeing Model for Model-Based Reinforcement Learning

Zifan Wu, Chao Yu, Chen Chen +2

In Model-based Reinforcement Learning (MBRL), model learning is critical since an inaccurate model can bias policy learning via generating misleading samples. However, learning an…