BEV-Seg: Bird's Eye View Semantic Segmentation Using Geometry and Semantic Point Cloud
arXiv:2006.11436
Abstract
Bird's-eye-view (BEV) is a powerful and widely adopted representation for road scenes that captures surrounding objects and their spatial locations, along with overall context in the scene. In this work, we focus on bird's eye semantic segmentation, a task that predicts pixel-wise semantic segmentation in BEV from side RGB images. This task is made possible by simulators such as Carla, which allow for cheap data collection, arbitrary camera placements, and supervision in ways otherwise not possible in the real world. There are two main challenges to this task: the view transformation from side view to bird's eye view, as well as transfer learning to unseen domains. Existing work transforms between views through fully connected layers and transfer learns via GANs. This suffers from a lack of depth reasoning and performance degradation across domains. Our novel 2-staged perception pipeline explicitly predicts pixel depths and combines them with pixel semantics in an efficient manner, allowing the model to leverage depth information to infer objects' spatial locations in the BEV. In addition, we transfer learning by abstracting high-level geometric features and predicting an intermediate representation that is common across different domains. We publish a new dataset called BEVSEG-Carla and show that our approach improves state-of-the-art by 24% mIoU and performs well when transferred to a new domain.
Accepted into CVPR 2020 Workshop Scalability in Autonomous Driving by Waymo
References in corpus (13)
- Unsupervised Cross-Domain Image Generation
- Deep High-Resolution Representation Learning for Visual Recognition
- IntentNet: Learning to Predict Intention from Raw Sensor Data
- Driving Policy Transfer via Modularity and Abstraction
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous Driving
- Beyond Grand Theft Auto V for Training, Testing and Enhancing Deep Learning in Self Driving Cars
- 3D Object Proposals using Stereo Imagery for Accurate Object Class Detection
- RefinedMPL: Refined Monocular PseudoLiDAR for 3D Object Detection in Autonomous Driving
- Minimal-Entropy Correlation Alignment for Unsupervised Deep Domain Adaptation
- A Geometric Approach to Obtain a Bird's Eye View from an Image
- Generative Adversarial Frontal View to Bird View Synthesis
- Deep Domain Adaptation by Geodesic Distance Minimization
- Pano2CAD: Room Layout From A Single Panorama Image
Cited by in corpus (5)
- Bird's-Eye-View Panoptic Segmentation Using Monocular Frontal View Images
- FIERY: Future Instance Prediction in Bird's-Eye View from Surround Monocular Cameras
- LaRa: Latents and Rays for Multi-Camera Bird's-Eye-View Semantic Segmentation
- Unlocking Past Information: Temporal Embeddings in Cooperative Bird's Eye View Prediction
- S-BEV: Semantic Birds-Eye View Representation for Weather and Lighting Invariant 3-DoF Localization