Zero-Shot Autonomous Vehicle Policy Transfer: From Simulation to Real-World via Adversarial Learning
arXiv:1903.05252 · doi:10.1109/ICCA51439.2020.9264552
Abstract
In this article, we demonstrate a zero-shot transfer of an autonomous driving policy from simulation to University of Delaware's scaled smart city with adversarial multi-agent reinforcement learning, in which an adversary attempts to decrease the net reward by perturbing both the inputs and outputs of the autonomous vehicles during training. We train the autonomous vehicles to coordinate with each other while crossing a roundabout in the presence of an adversary in simulation. The adversarial policy successfully reproduces the simulated behavior and incidentally outperforms, in terms of travel time, both a human-driving baseline and adversary-free trained policies. Finally, we demonstrate that the addition of adversarial training considerably improves the performance \eat{stability and robustness} of the policies after transfer to the real world compared to Gaussian noise injection.
6 pages, 4 figures
References in corpus (3)
Cited by in corpus (5)
- A Research and Educational Robotic Testbed for Real-time Control of Emerging Mobility Systems: From Theory to Scaled Experiments
- Adversarial Deep Reinforcement Learning for Improving the Robustness of Multi-agent Autonomous Driving Policies
- Combined Optimal Routing and Coordination of Connected and Automated Vehicles
- Stochastic Time-Optimal Trajectory Planning for Connected and Automated Vehicles in Mixed-Traffic Merging Scenarios
- A Hysteretic Q-learning Coordination Framework for Emerging Mobility Systems in Smart Cities