Learning Video Object Segmentation with Visual Memory
arXiv:1704.05737
Abstract
This paper addresses the task of segmenting moving objects in unconstrained videos. We introduce a novel two-stream neural network with an explicit memory module to achieve this. The two streams of the network encode spatial and temporal features in a video sequence respectively, while the memory module captures the evolution of objects over time. The module to build a "visual memory" in video, i.e., a joint representation of all the video frames, is realized with a convolutional recurrent unit learned from a small number of training video sequences. Given a video frame as input, our approach assigns each pixel an object or background label based on the learned spatio-temporal features as well as the "visual memory" specific to the video, acquired automatically without any manually-annotated frames. The visual memory is implemented with convolutional gated recurrent units, which allows to propagate spatial information over time. We evaluate our method extensively on two benchmarks, DAVIS and Freiburg-Berkeley motion segmentation datasets, and show state-of-the-art results. For example, our approach outperforms the top method on the DAVIS dataset by nearly 6%. We also provide an extensive ablative analysis to investigate the influence of each component in the proposed framework.
References in corpus (4)
Cited by in corpus (26)
- YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
- The 2018 DAVIS Challenge on Video Object Segmentation
- Video Object Segmentation and Tracking: A Survey
- Efficient Video Object Segmentation via Network Modulation
- Learning Inter- and Intraframe Representations for Non-Lambertian Photometric Stereo
- YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
- Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Adaptive Masked Proxies for Few-Shot Segmentation
- Interactive Medical Image Segmentation via Point-Based Interaction and Sequential Patch Learning
- Object Discovery in Videos as Foreground Motion Clustering
- Value of Temporal Dynamics Information in Driving Scene Segmentation
- Making a Case for 3D Convolutions for Object Segmentation in Videos
- See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks
- Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
- Visual Sensation and Perception Computational Models for Deep Learning: State of the art, Challenges and Prospects
- VideoMatch: Matching based Video Object Segmentation
- Spacetime Graph Optimization for Video Object Segmentation
- Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks
- Adaptive Temporal Encoding Network for Video Instance-level Human Parsing
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation
- Unsupervised Video Object Segmentation with Distractor-Aware Online Adaptation
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
- Unsupervised Video Object Segmentation using Motion Saliency-Guided Spatio-Temporal Propagation
- Non-parametric Memory for Spatio-Temporal Segmentation of Construction Zones for Self-Driving