Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes
arXiv:2101.06085
Abstract
Semantic segmentation is a key technology for autonomous vehicles to understand the surrounding scenes. The appealing performances of contemporary models usually come at the expense of heavy computations and lengthy inference time, which is intolerable for self-driving. Using light-weight architectures (encoder-decoder or two-pathway) or reasoning on low-resolution images, recent methods realize very fast scene parsing, even running at more than 100 FPS on a single 1080Ti GPU. However, there is still a significant gap in performance between these real-time methods and the models based on dilation backbones. To tackle this problem, we proposed a family of efficient backbones specially designed for real-time semantic segmentation. The proposed deep dual-resolution networks (DDRNets) are composed of two deep branches between which multiple bilateral fusions are performed. Additionally, we design a new contextual information extractor named Deep Aggregation Pyramid Pooling Module (DAPPM) to enlarge effective receptive fields and fuse multi-scale context based on low-resolution feature maps. Our method achieves a new state-of-the-art trade-off between accuracy and speed on both Cityscapes and CamVid dataset. In particular, on a single 2080Ti GPU, DDRNet-23-slim yields 77.4% mIoU at 102 FPS on Cityscapes test set and 74.7% mIoU at 230 FPS on CamVid test set. With widely used test augmentation, our method is superior to most state-of-the-art models and requires much less computation. Codes and trained models are available online.
12 pages, 7 figures. This work has been submitted to the IEEE for possible publication
References in corpus (7)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- High-Resolution Representations for Labeling Pixels and Regions
- Fast-SCNN: Fast Semantic Segmentation Network
- BiSeNet V2: Bilateral Network with Guided Aggregation for Real-time Semantic Segmentation
- FasterSeg: Searching for Faster Real-time Semantic Segmentation
- Real-Time Semantic Segmentation via Multiply Spatial Fusion Network
- Real-time Semantic Segmentation with Fast Attention
Cited by in corpus (18)
- RetiFluidNet: A Self-Adaptive and Multi-Attention Deep Convolutional Network for Retinal OCT Fluid Segmentation
- On the Real-World Adversarial Robustness of Real-Time Semantic Segmentation Models for Autonomous Driving
- A Comprehensive Review of Modern Object Segmentation Approaches
- SM-Net: Joint Learning of Semantic Segmentation and Stereo Matching for Autonomous Driving
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- P2AT: Pyramid Pooling Axial Transformer for Real-time Semantic Segmentation
- Defending From Physically-Realizable Adversarial Attacks Through Internal Over-Activation Analysis
- MMOCR: A Comprehensive Toolbox for Text Detection, Recognition and Understanding
- Rethinking Dilated Convolution for Real-time Semantic Segmentation
- Detection-segmentation convolutional neural network for autonomous vehicle perception
- Realtime Global Attention Network for Semantic Segmentation
- Attention-Based Real-Time Defenses for Physical Adversarial Attacks in Vision Applications
- Boundary Corrected Multi-scale Fusion Network for Real-time Semantic Segmentation
- Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object Detection
- CamLiFlow: Bidirectional Camera-LiDAR Fusion for Joint Optical Flow and Scene Flow Estimation
- Exploring the Effects of Data Augmentation for Drivable Area Segmentation
- Energy Consumption Analysis of pruned Semantic Segmentation Networks on an Embedded GPU
- Aerial-PASS: Panoramic Annular Scene Segmentation in Drone Videos