End-to-End Video Instance Segmentation with Transformers
arXiv:2011.14503
Abstract
Video instance segmentation (VIS) is the task that requires simultaneously classifying, segmenting and tracking object instances of interest in video. Recent methods typically develop sophisticated pipelines to tackle this task. Here, we propose a new video instance segmentation framework built upon Transformers, termed VisTR, which views the VIS task as a direct end-to-end parallel sequence decoding/prediction problem. Given a video clip consisting of multiple image frames as input, VisTR outputs the sequence of masks for each instance in the video in order directly. At the core is a new, effective instance sequence matching and segmentation strategy, which supervises and segments instances at the sequence level as a whole. VisTR frames the instance segmentation and tracking in the same perspective of similarity learning, thus considerably simplifying the overall pipeline and is significantly different from existing approaches. Without bells and whistles, VisTR achieves the highest speed among all existing VIS models, and achieves the best result among methods using single model on the YouTube-VIS dataset. For the first time, we demonstrate a much simpler and faster video instance segmentation framework built upon Transformers, achieving competitive accuracy. We hope that VisTR can motivate future research for more video understanding tasks.
CVPR2021 Oral
References in corpus (3)
Cited by in corpus (33)
- Transformers in Vision: A Survey
- Conditional Positional Encodings for Vision Transformers
- TransTrack: Multiple Object Tracking with Transformer
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- CvT: Introducing Convolutions to Vision Transformers
- All Tokens Matter: Token Labeling for Training Better Vision Transformers
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Efficient Self-supervised Vision Transformers for Representation Learning
- ASFormer: Transformer for Action Segmentation
- Video Instance Segmentation using Inter-Frame Communication Transformers
- TFPose: Direct Human Pose Estimation with Transformers
- CoTr: Efficiently Bridging CNN and Transformer for 3D Medical Image Segmentation
- Multiscale Vision Transformers
- Vision Transformer Pruning
- Refiner: Refining Self-attention for Vision Transformers
- TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
- A Video Is Worth Three Views: Trigeminal Transformers for Video-based Person Re-identification
- Segmenting Transparent Object in the Wild with Transformer
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- TransVOS: Video Object Segmentation with Transformers
- CoSformer: Detecting Co-Salient Object with Transformers
- TransLoc3D : Point Cloud based Large-scale Place Recognition using Adaptive Receptive Fields
- Delving Deep into the Generalization of Vision Transformers under Distribution Shifts
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
- End-to-End Video Object Detection with Spatial-Temporal Transformers
- Tracking Instances as Queries
- Sampling Equivariant Self-attention Networks for Object Detection in Aerial Images
- Crossover Learning for Fast Online Video Instance Segmentation
- Long-Short Temporal Contrastive Learning of Video Transformers
- Token Shift Transformer for Video Classification
- Ripple Attention for Visual Perception with Sub-quadratic Complexity