Joint 3D Proposal Generation and Object Detection from View Aggregation
arXiv:1712.02294
Abstract
We present AVOD, an Aggregate View Object Detection network for autonomous driving scenarios. The proposed neural network architecture uses LIDAR point clouds and RGB images to generate features that are shared by two subnetworks: a region proposal network (RPN) and a second stage detector network. The proposed RPN uses a novel architecture capable of performing multimodal feature fusion on high resolution feature maps to generate reliable 3D object proposals for multiple object classes in road scenes. Using these proposals, the second stage detection network performs accurate oriented 3D bounding box regression and category classification to predict the extents, orientation, and classification of objects in 3D space. Our proposed architecture is shown to produce state of the art results on the KITTI 3D object detection benchmark while running in real time with a low memory footprint, making it a suitable candidate for deployment on autonomous vehicles. Code is at: https://github.com/kujason/avod
For any inquiries contact aharakeh(at)uwaterloo(dot)ca
References in corpus (3)
Cited by in corpus (25)
- PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- 3D Object Detection Using Scale Invariant and Feature Reweighting Networks
- Orthographic Feature Transform for Monocular 3D Object Detection
- 3DSSD: Point-based 3D Single Stage Object Detector
- Complex-YOLO: Real-time 3D Object Detection on Point Clouds
- STD: Sparse-to-Dense 3D Object Detector for Point Cloud
- Deep SCNN-based Real-time Object Detection for Self-driving Vehicles Using LiDAR Temporal Data
- Using Machine Learning Safely in Automotive Software: An Assessment and Adaption of Software Process Requirements in ISO 26262
- Stereo R-CNN based 3D Object Detection for Autonomous Driving
- In Defense of Classical Image Processing: Fast Depth Completion on the CPU
- SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Robust Deep Multi-modal Learning Based on Gated Information Fusion Network
- Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection
- LMNet: Real-time Multiclass Object Detection on CPU using 3D LiDAR
- Multi-Sensor 3D Object Box Refinement for Autonomous Driving
- Deep Active Learning for Efficient Training of a LiDAR 3D Object Detector
- View Invariant Human Body Detection and Pose Estimation from Multiple Depth Sensors
- Class-specific Anchoring Proposal for 3D Object Recognition in LIDAR and RGB Images
- Joint Spatial-Temporal Optimization for Stereo 3D Object Tracking
- CubifAE-3D: Monocular Camera Space Cubification for Auto-Encoder based 3D Object Detection
- Leveraging Pre-Trained 3D Object Detection Models For Fast Ground Truth Generation
- Multi-scale Receptive Fields Graph Attention Network for Point Cloud Classification
- Know Your Surroundings: Panoramic Multi-Object Tracking by Multimodality Collaboration