activity
20182022
most citedLearning to Utilize Shaping Rewards: A New Approach of Reward Shaping

94 citations · 140 across the 8 of their papers we have counts for

collaborators

14 papers

cs.LG2022

A2C is a special case of PPO

Shengyi Huang, Anssi Kanervisto, Antonin Raffin +3

Advantage Actor-critic (A2C) and Proximal Policy Optimization (PPO) are popular deep reinforcement learning algorithms used for game AI in recent years. A common understanding is t…

cs.LG20221 cited

Coach-assisted Multi-Agent Reinforcement Learning Framework for Unexpected Crashed Agents

Jian Zhao, Youpeng Zhao, Weixun Wang +5

Multi-agent reinforcement learning is difficult to be applied in practice, which is partially due to the gap between the simulated and real-world scenarios. One reason for the gap…

cs.AI20223 cited

Revisiting QMIX: Discriminative Credit Assignment by Gradient Entropy Regularization

Jian Zhao, Yue Zhang, Xunhan Hu +5

In cooperative multi-agent systems, agents jointly take actions and receive a team reward instead of individual rewards. In the absence of individual reward signals, credit assignm…

cs.AI202110 cited

Cooperative Multi-Agent Transfer Learning with Level-Adaptive Credit Assignment

Tianze Zhou, Fubiao Zhang, Kun Shao +10

Extending transfer learning to cooperative multi-agent reinforcement learning (MARL) has recently received much attention. In contrast to the single-agent setting, the coordination…

cs.LG202094 cited

Learning to Utilize Shaping Rewards: A New Approach of Reward Shaping

Yujing Hu, Weixun Wang, Hangtian Jia +5

Reward shaping is an effective technique for incorporating domain knowledge into reinforcement learning (RL). Existing approaches such as potential-based reward shaping normally ma…

cs.DC2020

Learning to Accelerate Heuristic Searching for Large-Scale Maximum Weighted b-Matching Problems in Online Advertising

Xiaotian Hao, Junqi Jin, Jianye Hao +7

Bipartite b-matching is fundamental in algorithm design, and has been widely applied into economic markets, labor markets, etc. These practical problems usually exhibit two distinc…