PixelNet: Towards a General Pixel-level Architecture
arXiv:1609.06694
Abstract
We explore architectures for general pixel-level prediction problems, from low-level edge detection to mid-level surface normal estimation to high-level semantic segmentation. Convolutional predictors, such as the fully-convolutional network (FCN), have achieved remarkable success by exploiting the spatial redundancy of neighboring pixels through convolutional processing. Though computationally efficient, we point out that such approaches are not statistically efficient during learning precisely because spatial redundancy limits the information learned from neighboring pixels. We demonstrate that (1) stratified sampling allows us to add diversity during batch updates and (2) sampled multi-scale features allow us to explore more nonlinear predictors (multiple fully-connected layers followed by ReLU) that improve overall accuracy. Finally, our objective is to show how a architecture can get performance better than (or comparable to) the architectures designed for a particular task. Interestingly, our single architecture produces state-of-the-art results for semantic segmentation on PASCAL-Context, surface normal estimation on NYUDv2 dataset, and edge detection on BSDS without contextual post-processing.
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- ParseNet: Looking Wider to See Better
- Fully Convolutional Networks for Semantic Segmentation
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Learning Deconvolution Network for Semantic Segmentation
- FlowNet: Learning Optical Flow with Convolutional Networks
- DenseBox: Unifying Landmark Localization with End to End Object Detection
- Holistically-Nested Edge Detection
- Pixel-wise Deep Learning for Contour Detection
- Recurrent Convolutional Neural Networks for Scene Parsing
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- Top-Down Learning for Structured Labeling with Convolutional Pseudoprior
- The Fast Bilateral Solver
Cited by in corpus (20)
- An Implementation of Faster RCNN with Study for Region Sampling
- Subsurface structure analysis using computational interpretation and learning: A visual signal processing perspective
- Looking at Outfit to Parse Clothing
- Semantic Segmentation with Labeling Uncertainty and Class Imbalance
- Exploiting saliency for object segmentation from image level labels
- Improving Fully Convolution Network for Semantic Segmentation
- ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
- Gated Feedback Refinement Network for Coarse-to-Fine Dense Semantic Image Labeling
- Effective Use of Dilated Convolutions for Segmenting Small Object Instances in Remote Sensing Imagery
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Fast Face-swap Using Convolutional Neural Networks
- Self-supervised Learning for Single View Depth and Surface Normal Estimation
- Artificial intelligence based prediction on lung cancer risk factors using deep learning
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Mixed context networks for semantic segmentation
- Hierarchical Transfer Convolutional Neural Networks for Image Classification
- Modelling the Scene Dependent Imaging in Cameras with a Deep Neural Network
- Predicting Ground-Level Scene Layout from Aerial Imagery
- MT: Multi-Perspective Feature Learning Network for Scene Text Detection
- Deep Semantics-Aware Photo Adjustment