Deep Multi-agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
arXiv:2105.05701
Abstract
On-ramp merging is a challenging task for autonomous vehicles (AVs), especially in mixed traffic where AVs coexist with human-driven vehicles (HDVs). In this paper, we formulate the mixed-traffic highway on-ramp merging problem as a multi-agent reinforcement learning (MARL) problem, where the AVs (on both merge lane and through lane) collaboratively learn a policy to adapt to HDVs to maximize the traffic throughput. We develop an efficient and scalable MARL framework that can be used in dynamic traffic where the communication topology could be time-varying. Parameter sharing and local rewards are exploited to foster inter-agent cooperation while achieving great scalability. An action masking scheme is employed to improve learning efficiency by filtering out invalid/unsafe actions at each step. In addition, a novel priority-based safety supervisor is developed to significantly reduce collision rate and greatly expedite the training process. A gym-like simulation environment is developed and open-sourced with three different levels of traffic densities. We exploit curriculum learning to efficiently learn harder tasks from trained models under simpler settings. Comprehensive experimental results show the proposed MARL framework consistently outperforms several state-of-the-art benchmarks.
15 figures
References in corpus (11)
- Dota 2 with Large Scale Deep Reinforcement Learning
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- A Closer Look at Invalid Action Masking in Policy Gradient Algorithms
- Resource Management in Wireless Networks via Multi-Agent Deep Reinforcement Learning
- A DRL-based Multiagent Cooperative Control Framework for CAV Networks: a Graphic Convolution Q Network
- Safe Multi-Agent Reinforcement Learning via Shielding
- Leveraging the Capabilities of Connected and Autonomous Vehicles and Multi-Agent Reinforcement Learning to Mitigate Highway Bottleneck Congestion
- Safe Multi-Agent Reinforcement Learning through Decentralized Multiple Control Barrier Functions
- Revisiting Parameter Sharing in Multi-Agent Deep Reinforcement Learning
- PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control
- Combining Reinforcement Learning with Model Predictive Control for On-Ramp Merging