STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
arXiv:2003.08429 · doi:10.1007/978-3-030-58621-8_10
Abstract
Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by-detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in individual frames, and then associate these detections over time. Hence, these methods are often non-end-to-end trainable and highly tailored to specific tasks. In this paper, we propose a different approach that is well-suited to a variety of tasks involving instance segmentation in videos. In particular, we model a video clip as a single 3D spatio-temporal volume, and propose a novel approach that segments and tracks instances across space and time in a single stage. Our problem formulation is centered around the idea of spatio-temporal embeddings which are trained to cluster pixels belonging to a specific object instance over an entire video clip. To this end, we introduce (i) novel mixing functions that enhance the feature representation of spatio-temporal embeddings, and (ii) a single-stage, proposal-free network that can reason about temporal context. Our network is trained end-to-end to learn spatio-temporal embeddings as well as parameters required to cluster these embeddings, thus simplifying inference. Our method achieves state-of-the-art results across multiple datasets and tasks. Code and models are available at https://github.com/sabarim/STEm-Seg.
ECCV 2020 28 pages, 6 figures
References in corpus (8)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking
- Semantic Instance Segmentation with a Discriminative Loss Function
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation
- An Efficient 3D CNN for Action/Object Segmentation in Video
- Learning a Spatio-Temporal Embedding for Video Instance Segmentation
- Spatial Semantic Embedding Network: Fast 3D Instance Segmentation with Deep Metric Learning
Cited by in corpus (6)
- Solve the Puzzle of Instance Segmentation in Videos: A Weakly Supervised Framework with Spatio-Temporal Collaboration
- End-to-End Video Instance Segmentation with Transformers
- Video Instance Segmentation using Inter-Frame Communication Transformers
- Do Different Tracking Tasks Require Different Appearance Models?
- Weakly Supervised Instance Segmentation for Videos with Temporal Mask Consistency
- Opening up Open-World Tracking