Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning
arXiv:2001.08726
Abstract
Unlike popular modularized framework, end-to-end autonomous driving seeks to solve the perception, decision and control problems in an integrated way, which can be more adapting to new scenarios and easier to generalize at scale. However, existing end-to-end approaches are often lack of interpretability, and can only deal with simple driving tasks like lane keeping. In this paper, we propose an interpretable deep reinforcement learning method for end-to-end autonomous driving, which is able to handle complex urban scenarios. A sequential latent environment model is introduced and learned jointly with the reinforcement learning process. With this latent model, a semantic birdeye mask can be generated, which is enforced to connect with a certain intermediate property in today's modularized framework for the purpose of explaining the behaviors of learned policy. The latent space also significantly reduces the sample complexity of reinforcement learning. Comparison tests with a simulated autonomous car in CARLA show that the performance of our method in urban scenarios with crowded surrounding vehicles dominates many baselines including DQN, DDPG, TD3 and SAC. Moreover, through masked outputs, the learned policy is able to provide a better explanation of how the car reasons about the driving environment. The codes and videos of this work are available at our github repo and project website.
References in corpus (11)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic Algorithms and Applications
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- CARLA: An Open Urban Driving Simulator
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
- Deep Imitation Learning for Autonomous Driving in Generic Urban Scenarios with Enhanced Safety
- If MaxEnt RL is the Answer, What is the Question?
- PlaNet of the Bayesians: Reconsidering and Improving Deep Planning Network by Incorporating Bayesian Inference
Cited by in corpus (4)
- Task-Motion Planning for Safe and Efficient Urban Driving
- End-to-end Autonomous Driving Perception with Sequential Latent Representation Learning
- Goal-constrained Sparse Reinforcement Learning for End-to-End Driving
- A Safe Hierarchical Planning Framework for Complex Driving Scenarios based on Reinforcement Learning