TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
arXiv:2110.04507
Abstract
Deep reinforcement learning (DRL) has achieved super-human performance on complex video games (e.g., StarCraft II and Dota II). However, current DRL systems still suffer from challenges of multi-agent coordination, sparse rewards, stochastic environments, etc. In seeking to address these challenges, we employ a football video game, e.g., Google Research Football (GRF), as our testbed and develop an end-to-end learning-based AI system (denoted as TiKick) to complete this challenging task. In this work, we first generated a large replay dataset from the self-playing of single-agent experts, which are obtained from league training. We then developed a distributed learning system and new offline algorithms to learn a powerful multi-agent AI from the fixed single-agent dataset. To the best of our knowledge, Tikick is the first learning-based AI system that can take over the multi-agent Google Research Football full game, while previous work could either control a single agent or experiment on toy academic scenarios. Extensive experiments further show that our pre-trained model can accelerate the training process of the modern multi-agent algorithm and our method achieves state-of-the-art performances on various academic scenarios.
References in corpus (12)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Conservative Q-Learning for Offline Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Open-Ended Learning Leads to Generally Capable Agents
- Supervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of Kings
- Celebrating Diversity in Shared Multi-Agent Reinforcement Learning
- NeoRL: A Near Real-World Benchmark for Offline Reinforcement Learning
- On Bonus-Based Exploration Methods in the Arcade Learning Environment
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Simplifying Deep Reinforcement Learning via Self-Supervision
- Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum Games