FlowNet: Learning Optical Flow with Convolutional Networks
arXiv:1504.06852
Abstract
Convolutional neural networks (CNNs) have recently been very successful in a variety of computer vision tasks, especially on those linked to recognition. Optical flow estimation has not been among the tasks where CNNs were successful. In this paper we construct appropriate CNNs which are capable of solving the optical flow estimation problem as a supervised learning task. We propose and compare two architectures: a generic architecture and another one including a layer that correlates feature vectors at different image locations. Since existing ground truth data sets are not sufficiently large to train a CNN, we generate a synthetic Flying Chairs dataset. We show that networks trained on this unrealistic data still generalize very well to existing datasets such as Sintel and KITTI, achieving competitive accuracy at frame rates of 5 to 10 fps.
Added supplementary material
References in corpus (6)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Computing the Stereo Matching Cost with a Convolutional Neural Network
- Descriptor Matching with Convolutional Neural Networks: a Comparison to SIFT
- EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow
- Unsupervised learning of depth and motion
Cited by in corpus (116)
- Video Enhancement with Task-Oriented Flow
- Semantic Foggy Scene Understanding with Synthetic Data
- A Deep Learning Framework for Unsupervised Affine and Deformable Image Registration
- Weakly-Supervised Convolutional Neural Networks for Multimodal Image Registration
- Beyond Skip Connections: Top-Down Modulation for Object Detection
- Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
- Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM
- Slice-to-volume medical image registration: a survey
- Origami: A 803 GOp/s/W Convolutional Network Accelerator
- Neural SLAM: Learning to Explore with External Memory
- Super-Resolution via Deep Learning
- Look Wider to Match Image Patches with Convolutional Neural Networks
- Siamese Attentional Keypoint Network for High Performance Visual Tracking
- Hidden Two-Stream Convolutional Networks for Action Recognition
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Real-Time Deep Learning Approach to Visual Servo Control and Grasp Detection for Autonomous Robotic Manipulation
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- Learning Conditional Deformable Templates with Convolutional Networks
- Extracting full-field subpixel structural displacements from videos via deep learning
- PixelNet: Towards a General Pixel-level Architecture
- STFCN: Spatio-Temporal FCN for Semantic Video Segmentation
- DeepVO: A Deep Learning approach for Monocular Visual Odometry
- Learning Non-Metric Visual Similarity for Image Retrieval
- Deep Video Deblurring
- Coherent Online Video Style Transfer
- Procedural Modeling and Physically Based Rendering for Synthetic Data Generation in Automotive Applications
- Guided Optical Flow Learning
- Comparing recurrent and convolutional neural networks for predicting wave propagation
- Deep Multi-scale Convolutional Neural Network for Dynamic Scene Deblurring
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- Dual Semantic Fusion Network for Video Object Detection
- Deep Image Spatial Transformation for Person Image Generation
- Learning a Variational Network for Reconstruction of Accelerated MRI Data
- Real-Time and Accurate Object Detection in Compressed Video by Long Short-term Feature Aggregation
- Fully-Convolutional Siamese Networks for Object Tracking
- ALIGNet: Partial-Shape Agnostic Alignment via Unsupervised Learning
- Characterizing and Improving Stability in Neural Style Transfer
- ACDnet: An action detection network for real-time edge computing based on flow-guided feature approximation and memory aggregation
- Model-driven Simulations for Deep Convolutional Neural Networks
- 3D Human Pose Estimation from a Single Image via Distance Matrix Regression
- Beyond Tracking: Selecting Memory and Refining Poses for Deep Visual Odometry
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Memory Warps for Learning Long-Term Online Video Representations
- AirSim Drone Racing Lab
- Robustness Meets Deep Learning: An End-to-End Hybrid Pipeline for Unsupervised Learning of Egomotion
- Sequential Adversarial Learning for Self-Supervised Deep Visual Odometry
- An Artificial Agent for Robust Image Registration
- Unsupervised learning of object frames by dense equivariant image labelling
- Unsupervised Pose Flow Learning for Pose Guided Synthesis
- Simulations for Validation of Vision Systems
- Anytime Stereo Image Depth Estimation on Mobile Devices
- Deep Neural Network Concepts for Background Subtraction: A Systematic Review and Comparative Evaluation
- Temporal Interpolation via Motion Field Prediction
- Consistent Video Depth Estimation
- ActionFlowNet: Learning Motion Representation for Action Recognition
- BoundarySqueeze: Image Segmentation as Boundary Squeezing
- Surveillance Video Parsing with Single Frame Supervision
- Deep Rigid Instance Scene Flow
- A Transductive Approach for Video Object Segmentation
- ReHAR: Robust and Efficient Human Activity Recognition
- Learning Correspondence from the Cycle-Consistency of Time
- Natural Image Matting via Guided Contextual Attention
- Generative Models as a Data Source for Multiview Representation Learning
- TransFlow: Unsupervised Motion Flow by Joint Geometric and Pixel-level Estimation
- Selective Sensor Fusion for Neural Visual-Inertial Odometry
- Challenges in Time-Stamp Aware Anomaly Detection in Traffic Videos
- SceneEDNet: A Deep Learning Approach for Scene Flow Estimation
- Deep Depth From Focus
- Data-Driven Modeling of Geometry-Adaptive Steady Heat Transfer based on Convolutional Neural Networks: Heat Conduction
- STN-Homography: estimate homography parameters directly
- Attention-SLAM: A Visual Monocular SLAM Learning from Human Gaze
- Road obstacles positional and dynamic features extraction combining object detection, stereo disparity maps and optical flow data
- Learning Feature Descriptors using Camera Pose Supervision
- Data-driven 6D Pose Tracking by Calibrating Image Residuals in Synthetic Domains
- Temporal Modulation Network for Controllable Space-Time Video Super-Resolution
- Self-Supervised Object-in-Gripper Segmentation from Robotic Motions
- Automatic Generation of Dense Non-rigid Optical Flow
- Deep3D: Fully Automatic 2D-to-3D Video Conversion with Deep Convolutional Neural Networks
- Controllable Top-down Feature Transformer
- Solar Event Tracking with Deep Regression Networks: A Proof of Concept Evaluation
- HDRVideo-GAN: Deep Generative HDR Video Reconstruction
- Low-Latency Video Semantic Segmentation
- Multiple Object Tracking with Motion and Appearance Cues
- SIGNet: Semantic Instance Aided Unsupervised 3D Geometry Perception
- PatchBatch: a Batch Augmented Loss for Optical Flow
- Learning Global Structure Consistency for Robust Object Tracking
- Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling
- Adaptive Temporal Encoding Network for Video Instance-level Human Parsing
- Temporal Spatial-Adaptive Interpolation with Deformable Refinement for Electron Microscopic Images
- Learning Optical Flow from a Few Matches
- DAN: A Deformation-Aware Network for Consecutive Biomedical Image Interpolation
- Deep Visual Odometry with Adaptive Memory
- Face Image Reflection Removal
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- On the Sensory Commutativity of Action Sequences for Embodied Agents
- Learning Image Matching by Simply Watching Video
- Enhancing Multi-Robot Perception via Learned Data Association
- Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
- Indoor GeoNet: Weakly Supervised Hybrid Learning for Depth and Pose Estimation
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- CrossNet: An End-to-end Reference-based Super Resolution Network using Cross-scale Warping
- Gun Source and Muzzle Head Detection
- Video-based Person Re-Identification using Gated Convolutional Recurrent Neural Networks
- Deep Space-Time Video Upsampling Networks
- Revisiting Optical Flow Estimation in 360 Videos
- PICCOLO: Point Cloud-Centric Omnidirectional Localization
- Long-Short Temporal Modeling for Efficient Action Recognition
- Transformer Guided Geometry Model for Flow-Based Unsupervised Visual Odometry
- Deep Two-View Structure-from-Motion Revisited
- Visual Descriptor Learning from Monocular Video
- Understanding Road Layout from Videos as a Whole
- Weeping and Gnashing of Teeth: Teaching Deep Learning in Image and Video Processing Classes
- Semi-Supervised Semantic Matching
- Large-Scale Mapping of Human Activity using Geo-Tagged Videos
- Fully Convolutional Neural Networks for Dynamic Object Detection in Grid Maps (Masters Thesis)
- Guided Feature Selection for Deep Visual Odometry