A Comprehensive Review of Modern Object Segmentation Approaches
arXiv:2301.07499 · doi:10.1561/0600000097
Abstract
Image segmentation is the task of associating pixels in an image with their respective object class labels. It has a wide range of applications in many industries including healthcare, transportation, robotics, fashion, home improvement, and tourism. Many deep learning-based approaches have been developed for image-level object recognition and pixel-level scene understanding-with the latter requiring a much denser annotation of scenes with a large set of objects. Extensions of image segmentation tasks include 3D and video segmentation, where units of voxels, point clouds, and video frames are classified into different objects. We use "Object Segmentation" to refer to the union of these segmentation tasks. In this monograph, we investigate both traditional and modern object segmentation approaches, comparing their strengths, weaknesses, and utilities. We examine in detail the wide range of deep learning-based segmentation techniques developed in recent years, provide a review of the widely used datasets and evaluation metrics, and discuss potential future research directions.
173 pages, 49 figures, published in Foundations and Trends in Computer Graphics and Vision on 10/4/22. Authors retain copyright
References in corpus (19)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Neural Architecture Search with Reinforcement Learning
- Fully Convolutional Networks for Semantic Segmentation
- Semantic Segmentation using Adversarial Networks
- Hierarchical Multi-Scale Attention for Semantic Segmentation
- Semantic Instance Segmentation via Deep Metric Learning
- K-Net: Towards Unified Image Segmentation
- DeeperLab: Single-Shot Image Parser
- Polarization-driven Semantic Segmentation via Efficient Attention-bridged Fusion
- HoME: a Household Multimodal Environment
- SOLQ: Segmenting Objects by Learning Queries
- Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
- Scaling Wide Residual Networks for Panoptic Segmentation
- Auto-Panoptic: Cooperative Multi-Component Architecture Search for Panoptic Segmentation
- PanoNet: Real-time Panoptic Segmentation through Position-Sensitive Feature Embedding
- Tracking Instances as Queries
- Unifying Instance and Panoptic Segmentation with Dynamic Rank-1 Convolutions