Light-Weight RefineNet for Real-Time Semantic Segmentation
arXiv:1810.03272
Abstract
We consider an important task of effective and efficient semantic image segmentation. In particular, we adapt a powerful semantic segmentation architecture, called RefineNet, into the more compact one, suitable even for tasks requiring real-time performance on high-resolution inputs. To this end, we identify computationally expensive blocks in the original setup, and propose two modifications aimed to decrease the number of parameters and floating point operations. By doing that, we achieve more than twofold model reduction, while keeping the performance levels almost intact. Our fastest model undergoes a significant speed-up boost from 20 FPS to 55 FPS on a generic GPU card on 512x512 inputs with solid 81.1% mean iou performance on the test set of PASCAL VOC, while our slowest model with 32 FPS (from original 17 FPS) shows 82.7% mean iou on the same dataset. Alternatively, we showcase that our approach is easily mixable with light-weight classification networks: we attain 79.2% mean iou on PASCAL VOC using a model that contains only 3.3M parameters and performs only 9.3B floating point operations.
Models are available here: https://github.com/drsleep/light-weight-refinenet, BMVC 2018
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- WaveNet: A Generative Model for Raw Audio
- Compressing Deep Convolutional Networks using Vector Quantization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Incremental Network Quantization: Towards Lossless CNNs with Low-Precision Weights
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Dilated Residual Networks
- Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
Cited by in corpus (19)
- LiteSeg: A Novel Lightweight ConvNet for Semantic Segmentation
- Crowd Counting and Density Estimation by Trellis Encoder-Decoder Network
- A Comprehensive Review of Modern Object Segmentation Approaches
- Deep Multi-Branch Aggregation Network for Real-Time Semantic Segmentation in Street Scenes
- Density-based clustering with fully-convolutional networks for crowd flow detection from drones
- Residual Pyramid Learning for Single-Shot Semantic Segmentation
- Real-Time High-Performance Semantic Image Segmentation of Urban Street Scenes
- Auxiliary Learning for Deep Multi-task Learning
- NAS-Count: Counting-by-Density with Neural Architecture Search
- Edge-Guided Occlusion Fading Reduction for a Light-Weighted Self-Supervised Monocular Depth Estimation
- ShelfNet for Fast Semantic Segmentation
- Hard Pixel Mining for Depth Privileged Semantic Segmentation
- A Deep Learning Approach to Grasping the Invisible
- CI-Net: Contextual Information for Joint Semantic Segmentation and Depth Estimation
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- Perception Framework through Real-Time Semantic Segmentation and Scene Recognition on a Wearable System for the Visually Impaired
- Attention-Guided Lightweight Network for Real-Time Segmentation of Robotic Surgical Instruments
- Real-Time Selfie Video Stabilization
- Be Your Own Best Competitor! Multi-Branched Adversarial Knowledge Transfer