ParseNet: Looking Wider to See Better
arXiv:1506.04579
Abstract
We present a technique for adding global context to deep convolutional networks for semantic segmentation. The approach is simple, using the average feature for a layer to augment the features at each location. In addition, we study several idiosyncrasies of training, significantly increasing the performance of baseline networks (e.g. from FCN). When we add our proposed global feature, and a technique for learning normalization parameters, accuracy increases consistently even over our improved versions of the baselines. Our proposed approach, ParseNet, achieves state-of-the-art performance on SiftFlow and PASCAL-Context with small additional computational cost over baselines, and near current state-of-the-art performance on PASCAL VOC 2012 semantic segmentation with a simple approach. Code is available at https://github.com/weiliu89/caffe/tree/fcn .
ICLR 2016 submission
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Conditional Random Fields as Recurrent Neural Networks
- Going Deeper with Convolutions
- Fully Convolutional Networks for Semantic Segmentation
- Object Detectors Emerge in Deep Scene CNNs
- Fully Connected Deep Structured Networks
- Feedforward semantic segmentation with zoom-out features
Cited by in corpus (242)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Attention Mechanisms in Computer Vision: A Survey
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- A Review on Deep Learning Techniques Applied to Semantic Segmentation
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Fully Convolutional Networks for Semantic Segmentation
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Twins: Revisiting the Design of Spatial Attention in Vision Transformers
- OCNet: Object Context Network for Scene Parsing
- Segmentation Transformer: Object-Contextual Representations for Semantic Segmentation
- Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
- Feature Pyramid Networks for Object Detection
- DenseBox: Unifying Landmark Localization with End to End Object Detection
- CSPNet: A New Backbone that can Enhance Learning Capability of CNN
- Path Aggregation Network for Instance Segmentation
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- CCNet: Criss-Cross Attention for Semantic Segmentation
- Beyond Skip Connections: Top-Down Modulation for Object Detection
- A Fully Convolutional Neural Network for Cardiac Segmentation in Short-Axis MRI
- Pyramid Scene Parsing Network
- Semantic Labeling in Very High Resolution Images via a Self-Cascaded Convolutional Neural Network
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Pyramid Attention Network for Semantic Segmentation
- FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation
- Convolutional neural networks automate detection for tracking of submicron scale particles in 2D and 3D
- An Attention-Fused Network for Semantic Segmentation of Very-High-Resolution Remote Sensing Imagery
- Land Cover Classification from Remote Sensing Images Based on Multi-Scale Fully Convolutional Network
- DeeperLab: Single-Shot Image Parser
- What-and-Where to Match: Deep Spatially Multiplicative Integration Networks for Person Re-identification
- Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation
- Context Encoding for Semantic Segmentation
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Multi-scale fully convolutional neural networks for histopathology image segmentation: from nuclear aberrations to the global tissue architecture
- Dual Graph Convolutional Network for Semantic Segmentation
- Single-Shot Refinement Neural Network for Object Detection
- LPRNet: License Plate Recognition via Deep Neural Networks
- Tiny-DSOD: Lightweight Object Detection for Resource-Restricted Usages
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- Learning Representations for Automatic Colorization
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Revisiting Shadow Detection: A New Benchmark Dataset for Complex World
- Learning a Discriminative Feature Network for Semantic Segmentation
- Towards Automatic Concept-based Explanations
- LiteSeg: A Novel Lightweight ConvNet for Semantic Segmentation
- Stacked Deconvolutional Network for Semantic Segmentation
- Asymmetric Non-local Neural Networks for Semantic Segmentation
- GeoNet++: Iterative Geometric Neural Network with Edge-Aware Refinement for Joint Depth and Surface Normal Estimation
- Weakly Supervised Learning of Instance Segmentation with Inter-pixel Relations
- Pyramid Stereo Matching Network
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- Fully Convolutional Adaptation Networks for Semantic Segmentation
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Multi-Scale Feature Fusion: Learning Better Semantic Segmentation for Road Pothole Detection
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- Improving Fast Segmentation With Teacher-student Learning
- PixelNet: Towards a General Pixel-level Architecture
- Gated CRF Loss for Weakly Supervised Semantic Image Segmentation
- Multi-Stage Progressive Image Restoration
- Detect Faces Efficiently: A Survey and Evaluations
- MGNet: Monocular Geometric Scene Understanding for Autonomous Driving
- Prior Guided Feature Enrichment Network for Few-Shot Segmentation
- Global Aggregation then Local Distribution in Fully Convolutional Networks
- Accurate Single Stage Detector Using Recurrent Rolling Convolution
- ACFNet: Attentional Class Feature Network for Semantic Segmentation
- CANet: Class-Agnostic Segmentation Networks with Iterative Refinement and Attentive Few-Shot Learning
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- A Comprehensive Review of Modern Object Segmentation Approaches
- Attention-guided Chained Context Aggregation for Semantic Segmentation
- Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation
- Suppress and Balance: A Simple Gated Network for Salient Object Detection
- Learning Depth with Convolutional Spatial Propagation Network
- Improving Semantic Segmentation via Video Propagation and Label Relaxation
- Improving Semantic Segmentation via Self-Training
- SPGNet: Semantic Prediction Guidance for Scene Parsing
- Kronecker Attention Networks
- Vortex Pooling: Improving Context Representation in Semantic Segmentation
- Predicting dark matter halo formation in N-body simulations with deep regression networks
- Spatial-temporal Conv-sequence Learning with Accident Encoding for Traffic Flow Prediction
- GFF: Gated Fully Fusion for Semantic Segmentation
- SFD: Single Shot Scale-invariant Face Detector
- Deep Texture-Aware Features for Camouflaged Object Detection
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Camouflaged Object Segmentation with Distraction Mining
- Scaling Wide Residual Networks for Panoptic Segmentation
- Boosting RGB-D Saliency Detection by Leveraging Unlabeled RGB Images
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
- Improving Fully Convolution Network for Semantic Segmentation
- Spatial Pyramid Based Graph Reasoning for Semantic Segmentation
- Large-Field Contextual Feature Learning for Glass Detection
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation
- Squeeze-and-Attention Networks for Semantic Segmentation
- Efficient Yet Deep Convolutional Neural Networks for Semantic Segmentation
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- A Novel Upsampling and Context Convolution for Image Semantic Segmentation
- Global Aggregation then Local Distribution for Scene Parsing
- Selectivity or Invariance: Boundary-aware Salient Object Detection
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Multi-Scale Iterative Refinement Network for RGB-D Salient Object Detection
- A Single Stream Network for Robust and Real-time RGB-D Salient Object Detection
- Joint Semantic Segmentation and Boundary Detection using Iterative Pyramid Contexts
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
- Adaptable Deformable Convolutions for Semantic Segmentation of Fisheye Images in Autonomous Driving Systems
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Structure-Aware Network for Lane Marker Extraction with Dynamic Vision Sensor
- Deep Projective 3D Semantic Segmentation
- Flat2Layout: Flat Representation for Estimating Layout of General Room Types
- Graph-Based Global Reasoning Networks
- Hyper-Convolution Networks for Biomedical Image Segmentation
- Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection
- Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
- Optimizer Benchmarking Needs to Account for Hyperparameter Tuning
- Actor-Centric Relation Network
- Efficient 3D Fully Convolutional Networks for Pulmonary Lobe Segmentation in CT Images
- Multi-Path Feedback Recurrent Neural Network for Scene Parsing
- Customizable Architecture Search for Semantic Segmentation
- Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
- Real-Time High-Performance Semantic Image Segmentation of Urban Street Scenes
- Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Tensor Low-Rank Reconstruction for Semantic Segmentation
- A Survey on Deep Learning Methods for Semantic Image Segmentation in Real-Time
- PCN: Part and Context Information for Pedestrian Detection with CNNs
- Attentional Ptycho-Tomography (APT) for three-dimensional nanoscale X-ray imaging with minimal data acquisition and computation time
- SegSort: Segmentation by Discriminative Sorting of Segments
- STEP: Segmenting and Tracking Every Pixel
- DeepPerimeter: Indoor Boundary Estimation from Posed Monocular Sequences
- Weaving Multi-scale Context for Single Shot Detector
- DSFD: Dual Shot Face Detector
- ScratchDet: Training Single-Shot Object Detectors from Scratch
- MDFN: Multi-Scale Deep Feature Learning Network for Object Detection
- Decoders Matter for Semantic Segmentation: Data-Dependent Decoding Enables Flexible Feature Aggregation
- Multi-Scale Feature Aggregation by Cross-Scale Pixel-to-Region Relation Operation for Semantic Segmentation
- AlignSeg: Feature-Aligned Segmentation Networks
- Deep Scene Text Detection with Connected Component Proposals
- PFENet++: Boosting Few-shot Semantic Segmentation with the Noise-filtered Context-aware Prior Mask
- Dense Recurrent Neural Networks for Scene Labeling
- RT-K-Net: Revisiting K-Net for Real-Time Panoptic Segmentation
- Global and Local Features through Gaussian Mixture Models on Image Semantic Segmentation
- Is Faster R-CNN Doing Well for Pedestrian Detection?
- Learning to Predict Context-adaptive Convolution for Semantic Segmentation
- Deep Domain-Adversarial Image Generation for Domain Generalisation
- Multi-hypothesis contextual modeling for semantic segmentation
- Error Correction for Dense Semantic Image Labeling
- An Underwater Image Semantic Segmentation Method Focusing on Boundaries and a Real Underwater Scene Semantic Segmentation Dataset
- Differentiable Meta-learning Model for Few-shot Semantic Segmentation
- Attentional Feature Fusion
- TW-SMNet: Deep Multitask Learning of Tele-Wide Stereo Matching
- All you need are a few pixels: semantic segmentation with PixelPick
- CTNet: Context-based Tandem Network for Semantic Segmentation
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Background Subtraction with Real-time Semantic Segmentation
- Re-distributing Biased Pseudo Labels for Semi-supervised Semantic Segmentation: A Baseline Investigation
- ObjectAug: Object-level Data Augmentation for Semantic Image Segmentation
- 3D-to-2D Distillation for Indoor Scene Parsing
- Empirical Study of Multi-Task Hourglass Model for Semantic Segmentation Task
- From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real Data
- Triply Supervised Decoder Networks for Joint Detection and Segmentation
- Scene Labeling using Gated Recurrent Units with Explicit Long Range Conditioning
- A Unified Efficient Pyramid Transformer for Semantic Segmentation
- Bidirectional Projection Network for Cross Dimension Scene Understanding
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
- Improving Semantic Segmentation via Dilated Affinity
- Neuron-level Selective Context Aggregation for Scene Segmentation
- Multi-scale prediction for robust hand detection and classification
- Unsupervised Classification of Street Architectures Based on InfoGAN
- Learning Fully Dense Neural Networks for Image Semantic Segmentation
- Deep-Learning Assisted High-Resolution Binocular Stereo Depth Reconstruction
- Exploiting Global and Local Attentions for Heavy Rain Removal on Single Images
- A^2-FPN: Attention Aggregation based Feature Pyramid Network for Instance Segmentation
- Object-aware Feature Aggregation for Video Object Detection
- Distance Guided Channel Weighting for Semantic Segmentation
- Exploring Reciprocal Attention for Salient Object Detection by Cooperative Learning
- Occlusion-shared and Feature-separated Network for Occlusion Relationship Reasoning
- Consensus Feature Network for Scene Parsing
- What and Where: A Context-based Recommendation System for Object Insertion
- ShelfNet for Fast Semantic Segmentation
- Image Labeling with Markov Random Fields and Conditional Random Fields
- Contrast-Oriented Deep Neural Networks for Salient Object Detection
- Detecting Heads using Feature Refine Net and Cascaded Multi-Scale Architecture
- Spatial Memory for Context Reasoning in Object Detection
- Semi-supervised Semantic Segmentation with Directional Context-aware Consistency
- Temporal Feature Warping for Video Shadow Detection
- Deep Learning of Unified Region, Edge, and Contour Models for Automated Image Segmentation
- Attention-based fusion of semantic boundary and non-boundary information to improve semantic segmentation
- Fully Convolutional Network for Melanoma Diagnostics
- End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
- Context-aware Padding for Semantic Segmentation
- Automatic Real-time Background Cut for Portrait Videos
- Adaptive Context Encoding Module for Semantic Segmentation
- Rethinking ResNets: Improved Stacking Strategies With High Order Schemes
- Distilling Pixel-Wise Feature Similarities for Semantic Segmentation
- ELASTIC: Improving CNNs with Dynamic Scaling Policies
- Cross-Image Region Mining with Region Prototypical Network for Weakly Supervised Segmentation
- STD2P: RGBD Semantic Segmentation Using Spatio-Temporal Data-Driven Pooling
- High-Order Paired-ASPP Networks for Semantic Segmenation
- Object Detection from Scratch with Deep Supervision
- Object Detection with Mask-based Feature Encoding
- Learning Multi-level Region Consistency with Dense Multi-label Networks for Semantic Segmentation
- Multi-Path Region-Based Convolutional Neural Network for Accurate Detection of Unconstrained "Hard Faces"
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- Boundary Guidance Hierarchical Network for Real-Time Tongue Segmentation
- Exploiting Non-Local Priors via Self-Convolution For Highly-Efficient Image Restoration
- Perception-and-Regulation Network for Salient Object Detection
- Framework-agnostic Semantically-aware Global Reasoning for Segmentation
- Semantic-Rearrangement-Based Multi-Level Alignment for Domain Generalized Segmentation
- 3D Guided Weakly Supervised Semantic Segmentation
- DCANet: Dense Context-Aware Network for Semantic Segmentation
- Context Prior for Scene Segmentation
- StuffNet: Using 'Stuff' to Improve Object Detection
- MSDU-net: A Multi-Scale Dilated U-net for Blur Detection
- Multi-scale Interactive Network for Salient Object Detection
- GINet: Graph Interaction Network for Scene Parsing
- Adaptive Context Network for Scene Parsing
- Learning Where to Look While Tracking Instruments in Robot-assisted Surgery
- Growing a Brain: Fine-Tuning by Increasing Model Capacity
- Collaborative Global-Local Networks for Memory-Efficient Segmentation of Ultra-High Resolution Images
- Beyond Single Stage Encoder-Decoder Networks: Deep Decoders for Semantic Image Segmentation
- Neural Supervised Domain Adaptation by Augmenting Pre-trained Models with Random Units
- Automated segmentaiton and classification of arterioles and venules using Cascading Dilated Convolutional Neural Networks
- MPG-Net: Multi-Prediction Guided Network for Segmentation of Retinal Layers in OCT Images
- Hierarchical Semantic Segmentation using Psychometric Learning
- Salient Object Detection with Purificatory Mechanism and Structural Similarity Loss
- BiCANet: Bi-directional Contextual Aggregating Network for Image Semantic Segmentation
- Differentiating Features for Scene Segmentation Based on Dedicated Attention Mechanisms
- Affinity Derivation and Graph Merge for Instance Segmentation
- FHEDN: A based on context modeling Feature Hierarchy Encoder-Decoder Network for face detection
- Gesture-based Bootstrapping for Egocentric Hand Segmentation
- Rethinking Fully Convolutional Networks for the Analysis of Photoluminescence Wafer Images
- Semantic Segmentation for Urban-Scene Images
- Sparse Spatial Attention Network for Semantic Segmentation
- MSFD:Multi-Scale Receptive Field Face Detector
- Boundary-induced and scene-aggregated network for monocular depth prediction
- Exploring Temporal Information for Improved Video Understanding
- Distortion-adaptive Salient Object Detection in 360 Omnidirectional Images
- An Abstraction Model for Semantic Segmentation Algorithms
- Superpixel-based Semantic Segmentation Trained by Statistical Process Control
- Attention to Refine through Multi-Scales for Semantic Segmentation
- CMF: Cascaded Multi-model Fusion for Referring Image Segmentation
- Chinese Herbal Recognition based on Competitive Attentional Fusion of Multi-hierarchies Pyramid Features