UVOSAM: A Mask-free Paradigm for Unsupervised Video Object Segmentation via Segment Anything Model
arXiv:2305.12659 · doi:10.1016/j.patcog.2024.111100
Abstract
The current state-of-the-art methods for unsupervised video object segmentation (UVOS) require extensive training on video datasets with mask annotations, limiting their effectiveness in handling challenging scenarios. However, the Segment Anything Model (SAM) introduces a new prompt-driven paradigm for image segmentation, offering new possibilities. In this study, we investigate SAM's potential for UVOS through different prompt strategies. We then propose UVOSAM, a mask-free paradigm for UVOS that utilizes the STD-Net tracker. STD-Net incorporates a spatial-temporal decoupled deformable attention mechanism to establish an effective correlation between intra- and inter-frame features, remarkably enhancing the quality of box prompts in complex video scenes. Extensive experiments on the DAVIS2017-unsupervised and YoutubeVIS19\&21 datasets demonstrate the superior performance of UVOSAM without mask supervision compared to existing mask-supervised methods, as well as its ability to generalize to weakly-annotated video datasets. Code can be found at https://github.com/alibaba/UVOSAM.
journal = {Pattern Recognition}
References in corpus (18)
- Representation Learning with Contrastive Predictive Coding
- YOLOX: Exceeding YOLO Series in 2021
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- MOT20: A benchmark for multi object tracking in crowded scenes
- TransTrack: Multiple Object Tracking with Transformer
- Segment Everything Everywhere All at Once
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object Segmentation
- Track Anything: Segment Anything Meets Videos
- Segment and Track Anything
- Mask2Former for Video Instance Segmentation
- SegGPT: Segmenting Everything In Context
- Video Instance Segmentation using Inter-Frame Communication Transformers
- Segment Anything Meets Point Tracking
- In-N-Out Generative Learning for Dense Unsupervised Video Segmentation
- Dual Prototype Attention for Unsupervised Video Object Segmentation
- Maximal Cliques on Multi-Frame Proposal Graph for Unsupervised Video Object Segmentation
- LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation