1 paper
Cuiling Wu, Yaozhong Gan, Junliang Xing +1
We propose Multi Agent Reflective Policy Optimization (MARPO) to alleviate the issue of sample inefficiency in multi agent reinforcement learning. MARPO consists of two key compone…