Recurrent Convolutional Neural Networks for Scene Parsing
arXiv:1306.2795
Abstract
Scene parsing is a technique that consist on giving a label to all pixels in an image according to the class they belong to. To ensure a good visual coherence and a high class accuracy, it is essential for a scene parser to capture image long range dependencies. In a feed-forward architecture, this can be simply achieved by considering a sufficiently large input context patch, around each pixel to be labeled. We propose an approach consisting of a recurrent convolutional neural network which allows us to consider a large input context, while limiting the capacity of the model. Contrary to most standard approaches, our method does not rely on any segmentation methods, nor any task-specific features. The system is trained in an end-to-end manner over raw pixels, and models complex spatial dependencies with low inference cost. As the context size increases with the built-in recurrence, the system identifies and corrects its own errors. Our approach yields state-of-the-art performance on both the Stanford Background Dataset and the SIFT Flow Dataset, while remaining very fast at test time.
Cited by in corpus (25)
- Perceptual Losses for Real-Time Style Transfer and Super-Resolution
- How to Construct Deep Recurrent Neural Networks
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- DISC: Deep Image Saliency Computing via Progressive Representation Learning
- Swapout: Learning an ensemble of deep architectures
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Efficient piecewise training of deep structured models for semantic segmentation
- A Survey of Semantic Segmentation
- PixelNet: Towards a General Pixel-level Architecture
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- rnn : Recurrent Library for Torch
- Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-expressions
- MUST-CNN: A Multilayer Shift-and-Stitch Deep Convolutional Architecture for Sequence-based Protein Structure Prediction
- Multi-Path Feedback Recurrent Neural Network for Scene Parsing
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
- Unsupervised Total Variation Loss for Semi-supervised Deep Learning of Semantic Segmentation
- Exploring Context with Deep Structured models for Semantic Segmentation
- Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
- Weakly Supervised Object Localization Using Things and Stuff Transfer
- Integrated perception with recurrent multi-task neural networks
- Image Segmentation for Fruit Detection and Yield Estimation in Apple Orchards
- Top-down Neural Attention by Excitation Backprop
- DeepChrome: Deep-learning for predicting gene expression from histone modifications
- Object Boundary Guided Semantic Segmentation
- Region-based semantic segmentation with end-to-end training