Learning Video Object Segmentation from Static Images
arXiv:1612.02646
Abstract
Inspired by recent advances of deep learning in instance segmentation and object tracking, we introduce video object segmentation problem as a concept of guided instance segmentation. Our model proceeds on a per-frame basis, guided by the output of the previous frame towards the object of interest in the next frame. We demonstrate that highly accurate object segmentation in videos can be enabled by using a convnet trained with static images only. The key ingredient of our approach is a combination of offline and online learning strategies, where the former serves to produce a refined mask from the previous frame estimate and the latter allows to capture the appearance of the specific object instance. Our method can handle different types of input annotations: bounding boxes and segments, as well as incorporate multiple annotated frames, making the system suitable for diverse applications. We obtain competitive results on three different datasets, independently from the type of input annotation.
Submitted to CVPR 2017
References in corpus (2)
Cited by in corpus (44)
- Video Salient Object Detection via Fully Convolutional Networks
- RANet: Ranking Attention Network for Fast Video Object Segmentation
- Fast Online Object Tracking and Segmentation: A Unifying Approach
- SegFlow: Joint Learning for Video Object Segmentation and Optical Flow
- Learning Video Object Segmentation with Visual Memory
- FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation
- Video Propagation Networks
- Lucid Data Dreaming for Video Object Segmentation
- Anchor Diffusion for Unsupervised Video Object Segmentation
- Full-Duplex Strategy for Video Object Segmentation
- Learning to Segment Instances in Videos with Spatial Propagation Network
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- Actor-Action Semantic Segmentation with Region Masks
- MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation
- Prediction-Tracking-Segmentation
- Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion
- DMM-Net: Differentiable Mask-Matching Network for Video Object Segmentation
- Efficient Regional Memory Network for Video Object Segmentation
- Fast Video Object Segmentation using the Global Context Module
- MAST: A Memory-Augmented Self-supervised Tracker
- Spatiotemporal CNN for Video Object Segmentation
- Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation
- Video Object Segmentation using Supervoxel-Based Gerrymandering
- Learning About Objects by Learning to Interact with Them
- Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation
- Low-Latency Video Semantic Segmentation
- See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks
- MSN: Efficient Online Mask Selection Network for Video Instance Segmentation
- OVSNet : Towards One-Pass Real-Time Video Object Segmentation
- SwiftNet: Real-time Video Object Segmentation
- Spacetime Graph Optimization for Video Object Segmentation
- Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template Matching
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- Automatic Real-time Background Cut for Portrait Videos
- Discriminative Online Learning for Fast Video Object Segmentation
- Video Panoptic Segmentation
- Fast Template Matching and Update for Video Object Tracking and Segmentation
- Learning Dynamic Network Using a Reuse Gate Function in Semi-supervised Video Object Segmentation
- Unseen Object Segmentation in Videos via Transferable Representations
- Fast Video Object Segmentation via Mask Transfer Network
- DAVOS: Semi-Supervised Video Object Segmentation via Adversarial Domain Adaptation
- Space Time Recurrent Memory Network
- CNN in MRF: Video Object Segmentation via Inference in A CNN-Based Higher-Order Spatio-Temporal MRF