activity
20182022
most citedLearning to Utilize Shaping Rewards: A New Approach of Reward Shaping

94 citations · 161 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2022★ 11 cited

MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning Library

Siyi Hu, Yifan Zhong, Minquan Gao +6

A significant challenge facing researchers in the area of multi-agent reinforcement learning (MARL) pertains to the identification of a library that can offer fast and compatible d…

cs.LG2022

A2C is a special case of PPO

Shengyi Huang, Anssi Kanervisto, Antonin Raffin +3

Advantage Actor-critic (A2C) and Proximal Policy Optimization (PPO) are popular deep reinforcement learning algorithms used for game AI in recent years. A common understanding is t…

cs.LG2022★ 1 cited

Coach-assisted Multi-Agent Reinforcement Learning Framework for Unexpected Crashed Agents

Jian Zhao, Youpeng Zhao, Weixun Wang +5

Multi-agent reinforcement learning is difficult to be applied in practice, which is partially due to the gap between the simulated and real-world scenarios. One reason for the gap…

cs.LG2022★ 10 cited

Breaking the Curse of Dimensionality in Multiagent State Space: A Unified Agent Permutation Framework

Xiaotian Hao, Hangyu Mao, Weixun Wang +5

The state space in Multiagent Reinforcement Learning (MARL) grows exponentially with the agent number. Such a curse of dimensionality results in poor scalability and low sample eff…

cs.LG2020★ 94 cited

Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping

Yujing Hu, Weixun Wang, Hangtian Jia +5

Reward shaping is an effective technique for incorporating domain knowledge into reinforcement learning (RL). Existing approaches such as potential-based reward shaping normally ma…

cs.LG2020

Efficient Deep Reinforcement Learning via Adaptive Policy Transfer

Tianpei Yang, Jianye Hao, Zhaopeng Meng +8

Transfer Learning (TL) has shown great potential to accelerate Reinforcement Learning (RL) by leveraging prior knowledge from past learned policies of relevant tasks. Existing tran…