Rethinking Closed-loop Training for Autonomous Driving
arXiv:2306.15713 · doi:10.1007/978-3-031-19842-7_16
Abstract
Recent advances in high-fidelity simulators have enabled closed-loop training of autonomous driving agents, potentially solving the distribution shift in training v.s. deployment and allowing training to be scaled both safely and cheaply. However, there is a lack of understanding of how to build effective training benchmarks for closed-loop training. In this work, we present the first empirical study which analyzes the effects of different training benchmark designs on the success of learning agents, such as how to design traffic scenarios and scale training environments. Furthermore, we show that many popular RL algorithms cannot achieve satisfactory performance in the context of autonomous driving, as they lack long-term planning and take an extremely long time to train. To address these issues, we propose trajectory value learning (TRAVL), an RL-based driving agent that performs planning with multistep look-ahead and exploits cheaply generated imagined data for efficient learning. Our experiments show that TRAVL can learn much faster and produce safer maneuvers compared to all the baselines. For more information, visit the project website: https://waabi.ai/research/travl
ECCV 2022
References in corpus (5)
- Benchmarking Model-Based Reinforcement Learning
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic
- A Survey on Scenario-Based Testing for Automated Driving Systems in High-Fidelity Simulation
- Offline Reinforcement Learning for Autonomous Driving with Safety and Exploration Enhancement