FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
arXiv:1701.05384
Abstract
We propose an end-to-end learning framework for segmenting generic objects in videos. Our method learns to combine appearance and motion information to produce pixel level segmentation masks for all prominent objects in videos. We formulate this task as a structured prediction problem and design a two-stream fully convolutional neural network which fuses together motion and appearance in a unified framework. Since large-scale video datasets with pixel level segmentations are problematic, we show how to bootstrap weakly annotated videos together with existing image recognition datasets for training. Through experiments on three challenging video segmentation benchmarks, our method substantially improves the state-of-the-art for segmenting generic (unseen) objects. Code and pre-trained models are available on the project website.
CVPR 2017
References in corpus (3)
Cited by in corpus (63)
- Anabranch Network for Camouflaged Object Segmentation
- YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
- The 2018 DAVIS Challenge on Video Object Segmentation
- STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos
- The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation
- River Ice Segmentation with Deep Learning
- Data augmentation using learned transformations for one-shot medical image segmentation
- MODNet: Moving Object Detection Network with Motion and Appearance for Autonomous Driving
- Video Object Segmentation and Tracking: A Survey
- RVOS: End-to-End Recurrent Network for Video Object Segmentation
- Lucid Data Dreaming for Video Object Segmentation
- An Efficient 3D CNN for Action/Object Segmentation in Video
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- Anchor Diffusion for Unsupervised Video Object Segmentation
- Full-Duplex Strategy for Video Object Segmentation
- YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
- Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning
- Motion Guided Attention for Video Salient Object Detection
- RST-MODNet: Real-time Spatio-temporal Moving Object Detection for Autonomous Driving
- Adaptive Masked Proxies for Few-Shot Segmentation
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Video Instance Segmentation
- MaskRNN: Instance Level Video Object Segmentation
- Actor-Action Semantic Segmentation with Region Masks
- Mutual Suppression Network for Video Prediction using Disentangled Features
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Fast and Robust Dynamic Hand Gesture Recognition via Key Frames Extraction and Feature Fusion
- ALBA : Reinforcement Learning for Video Object Segmentation
- Object Discovery in Videos as Foreground Motion Clustering
- Spatiotemporal CNN for Video Object Segmentation
- Self-supervised Video Object Segmentation by Motion Grouping
- Video Object Segmentation using Supervoxel-Based Gerrymandering
- TTVOS: Lightweight Video Object Segmentation with Adaptive Template Attention Module and Temporal Consistency Loss
- Multi-Source Fusion and Automatic Predictor Selection for Zero-Shot Video Object Segmentation
- Making a Case for 3D Convolutions for Object Segmentation in Videos
- Self-Supervised Object-in-Gripper Segmentation from Robotic Motions
- Learning Discriminative Feature with CRF for Unsupervised Video Object Segmentation
- See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks
- Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos
- One-Shot Weakly Supervised Video Object Segmentation
- Zero-Shot Video Object Segmentation via Attentive Graph Neural Networks
- Patchwork: A Patch-wise Attention Network for Efficient Object Detection and Segmentation in Video Streams
- SwiftNet: Real-time Video Object Segmentation
- VideoMatch: Matching based Video Object Segmentation
- Spacetime Graph Optimization for Video Object Segmentation
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- Self-supervised Training of Proposal-based Segmentation via Background Prediction
- Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation
- Fast Video Object Segmentation via Mask Transfer Network
- Unsupervised Video Object Segmentation with Distractor-Aware Online Adaptation
- Guided Interactive Video Object Segmentation Using Reliability-Based Attention Maps
- Learning Dynamic Network Using a Reuse Gate Function in Semi-supervised Video Object Segmentation
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object Segmentation
- Unseen Object Segmentation in Videos via Transferable Representations
- Flow Based Self-supervised Pixel Embedding for Image Segmentation
- DAVOS: Semi-Supervised Video Object Segmentation via Adversarial Domain Adaptation
- Beyond Single Stage Encoder-Decoder Networks: Deep Decoders for Semantic Image Segmentation
- Video Region Annotation with Sparse Bounding Boxes
- Unsupervised Video Object Segmentation using Motion Saliency-Guided Spatio-Temporal Propagation
- Automatic Video Object Segmentation via Motion-Appearance-Stream Fusion and Instance-aware Segmentation
- Online Mutual Foreground Segmentation for Multispectral Stereo Videos
- Learning to Segment Rigid Motions from Two Frames
- Learning a Weakly-Supervised Video Actor-Action Segmentation Model with a Wise Selection