Semantic Understanding of Scenes through the ADE20K Dataset
arXiv:1608.05442
Abstract
Scene parsing, or recognizing and segmenting objects and stuff in an image, is one of the key problems in computer vision. Despite the community's efforts in data collection, there are still few image datasets covering a wide range of scenes and object categories with dense and detailed annotations for scene parsing. In this paper, we introduce and analyze the ADE20K dataset, spanning diverse annotations of scenes, objects, parts of objects, and in some cases even parts of parts. A generic network design called Cascade Segmentation Module is then proposed to enable the segmentation networks to parse a scene into stuff, objects, and object parts in a cascade. We evaluate the proposed module integrated within two existing semantic segmentation networks, yielding significant improvements for scene parsing. We further show that the scene parsing networks trained on ADE20K can be applied to a wide variety of scenes and objects.
IJCV extension
References in corpus (6)
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Learning Deconvolution Network for Semantic Segmentation
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
- COCO-Stuff: Thing and Stuff Classes in Context
Cited by in corpus (45)
- Adversarial Learning for Semi-Supervised Semantic Segmentation
- Pyramid Scene Parsing Network
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Context Based Emotion Recognition using EMOTIC Dataset
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Places: An Image Database for Deep Scene Understanding
- SegICP: Integrated Deep Semantic Segmentation and Pose Estimation
- Learning to Generate Images of Outdoor Scenes from Attributes and Semantic Layouts
- Dual Graph Convolutional Network for Semantic Segmentation
- SESAME: Semantic Editing of Scenes by Adding, Manipulating or Erasing Objects
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- DFANet: Deep Feature Aggregation for Real-Time Semantic Segmentation
- Evidential fully convolutional network for semantic segmentation
- Visual Affordance and Function Understanding: A Survey
- Annotating Object Instances with a Polygon-RNN
- GFF: Gated Fully Fusion for Semantic Segmentation
- Explaining Trained Neural Networks with Semantic Web Technologies: First Steps
- Iterative Visual Reasoning Beyond Convolutions
- Improving Fully Convolution Network for Semantic Segmentation
- Global Aggregation then Local Distribution for Scene Parsing
- Deep Image Harmonization
- TextMountain: Accurate Scene Text Detection via Instance Segmentation
- Semantic Flow for Fast and Accurate Scene Parsing
- Semantic White Balance: Semantic Color Constancy Using Convolutional Neural Network
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Setting an attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels
- PPGNet: Learning Point-Pair Graph for Line Segment Detection
- PointFlow: Flowing Semantics Through Points for Aerial Image Segmentation
- FishEyeRecNet: A Multi-Context Collaborative Deep Network for Fisheye Image Rectification
- Dynamic-structured Semantic Propagation Network
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Hierarchical semantic segmentation using modular convolutional neural networks
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Deep Learning--Based Scene Simplification for Bionic Vision
- Mixed context networks for semantic segmentation
- Deep neural networks can be improved using human-derived contextual expectations
- Reducing the feature divergence of RGB and near-infrared images using Switchable Normalization
- The Unreasonable Effectiveness of Texture Transfer for Single Image Super-resolution
- Semantic Segmentation via Highly Fused Convolutional Network with Multiple Soft Cost Functions
- Efficient Convolutional Neural Network with Binary Quantization Layer
- Deep Image Synthesis from Intuitive User Input: A Review and Perspectives
- The CASE Dataset of Candidate Spaces for Advert Implantation
- Learning to Label Affordances from Simulated and Real Data
- MoE-SPNet: A Mixture-of-Experts Scene Parsing Network
- Attention to Refine through Multi-Scales for Semantic Segmentation