Simultaneous Detection and Segmentation
arXiv:1407.1808
Abstract
We aim to detect all instances of a category in an image and, for each instance, mark the pixels that belong to it. We call this task Simultaneous Detection and Segmentation (SDS). Unlike classical bounding box detection, SDS requires a segmentation and not just a box. Unlike classical semantic segmentation, we require individual object instances. We build on recent work that uses convolutional neural networks to classify category-independent region proposals (R-CNN [16]), introducing a novel architecture tailored for SDS. We then use category-specific, top- down figure-ground predictions to refine our bottom-up proposals. We show a 7 point boost (16% relative) over our baselines on SDS, a 5 point boost (10% relative) over state-of-the-art on semantic segmentation, and state-of-the-art performance in object detection. Finally, we provide diagnostic tools that unpack performance and provide directions for future work.
To appear in the European Conference on Computer Vision (ECCV), 2014
Cited by in corpus (15)
- Learning Deconvolution Network for Semantic Segmentation
- Semantic Instance Segmentation via Deep Metric Learning
- Deep convolutional filter banks for texture recognition and segmentation
- Inferring 3D Object Pose in RGB-D Images
- Convolutional Neural Networks at Constrained Time Cost
- Joint Object and Part Segmentation using Deep Learned Potentials
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Hypercolumns for Object Segmentation and Fine-grained Localization
- From Image-level to Pixel-level Labeling with Convolutional Networks
- Cultural Event Recognition with Visual ConvNets and Temporal Models
- Learning Rich Features from RGB-D Images for Object Detection and Segmentation
- Category-Specific Object Reconstruction from a Single Image
- Symbolic Segmentation Using Algorithm Selection
- Candidate Constrained CRFs for Loss-Aware Structured Prediction
- Subset Feature Learning for Fine-Grained Category Classification