COCO-Stuff: Thing and Stuff Classes in Context
arXiv:1612.03716
Abstract
Semantic classes can be either things (objects with a well-defined shape, e.g. car, person) or stuff (amorphous background regions, e.g. grass, sky). While lots of classification and detection works focus on thing classes, less attention has been given to stuff classes. Nonetheless, stuff classes are important as they allow to explain important aspects of an image, including (1) scene type; (2) which thing classes are likely to be present and their location (through contextual reasoning); (3) physical attributes, material types and geometric properties of the scene. To understand stuff and things in context we introduce COCO-Stuff, which augments all 164K images of the COCO 2017 dataset with pixel-wise annotations for 91 stuff classes. We introduce an efficient stuff annotation protocol based on superpixels, which leverages the original thing annotations. We quantify the speed versus quality trade-off of our protocol and explore the relation between annotation time and boundary complexity. Furthermore, we use COCO-Stuff to analyze: (a) the importance of stuff and thing classes in terms of their surface cover and how frequently they are mentioned in image captions; (b) the spatial relations between stuff and things, highlighting the rich contextual relations that make our dataset unique; (c) the performance of a modern semantic segmentation method on stuff and thing classes, and whether stuff is easier to segment than things.
CVPR 2018 camera-ready
References in corpus (6)
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Simultaneous Detection and Segmentation
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Annotating Object Instances with a Polygon-RNN
- Weakly Supervised Object Localization Using Things and Stuff Transfer
Cited by in corpus (19)
- Dual Attention Network for Scene Segmentation
- Hybrid Task Cascade for Instance Segmentation
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Interactive Image Generation Using Scene Graphs
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Boundary-Aware Network for Fast and High-Accuracy Portrait Segmentation
- Text2Scene: Generating Compositional Scenes from Textual Descriptions
- Setting an attention region for convolutional neural networks using region selective features, for recognition of materials within glass vessels
- Learning Instance Occlusion for Panoptic Segmentation
- Tree-structured Kronecker Convolutional Network for Semantic Segmentation
- Image Generation from Layout
- SketchyCOCO: Image Generation from Freehand Scene Sketches
- Dynamic-structured Semantic Propagation Network
- Value of Temporal Dynamics Information in Driving Scene Segmentation
- Hierarchical semantic segmentation using modular convolutional neural networks
- Locally Adaptive Learning Loss for Semantic Image Segmentation
- Boundary-sensitive Network for Portrait Segmentation
- Neuron-level Selective Context Aggregation for Scene Segmentation
- Visual Semantic Information Pursuit: A Survey