Unsupervised Video Object Segmentation with Distractor-Aware Online Adaptation
arXiv:1812.07712
Abstract
Unsupervised video object segmentation is a crucial application in video analysis without knowing any prior information about the objects. It becomes tremendously challenging when multiple objects occur and interact in a given video clip. In this paper, a novel unsupervised video object segmentation approach via distractor-aware online adaptation (DOA) is proposed. DOA models spatial-temporal consistency in video sequences by capturing background dependencies from adjacent frames. Instance proposals are generated by the instance segmentation network for each frame and then selected by motion information as hard negatives if they exist and positives. To adopt high-quality hard negatives, the block matching algorithm is then applied to preceding frames to track the associated hard negatives. General negatives are also introduced in case that there are no hard negatives in the sequence and experiments demonstrate both kinds of negatives (distractors) are complementary. Finally, we conduct DOA using the positive, negative, and hard negative masks to update the foreground/background segmentation. The proposed approach achieves state-of-the-art results on two benchmark datasets, DAVIS 2016 and FBMS-59 datasets.
11 pages, 6 figures, 4 tables, conference
References in corpus (15)
- Adam: A Method for Stochastic Optimization
- TensorFlow: A system for large-scale machine learning
- Fully Convolutional Networks for Semantic Segmentation
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Simultaneous Detection and Segmentation
- FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
- Bootstrapping Face Detection with Hard Negative Examples
- Learning Video Object Segmentation with Visual Memory
- Semantic Segmentation with Reverse Attention
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Unsupervised Hard Example Mining from Videos for Improved Object Detection
- VideoMatch: Matching based Video Object Segmentation
- Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation