SfM-Net: Learning of Structure and Motion from Video
arXiv:1704.07804
Abstract
We propose SfM-Net, a geometry-aware neural network for motion estimation in videos that decomposes frame-to-frame pixel motion in terms of scene and object depth, camera motion and 3D object rotations and translations. Given a sequence of frames, SfM-Net predicts depth, segmentation, camera and rigid object motions, converts those into a dense frame-to-frame motion field (optical flow), differentiably warps frames in time to match pixels and back-propagates. The model can be trained with various degrees of supervision: 1) self-supervised by the re-projection photometric error (completely unsupervised), 2) supervised by ego-motion (camera motion), or 3) supervised by depth (e.g., as provided by RGBD sensors). SfM-Net extracts meaningful depth estimates and successfully estimates frame-to-frame camera rotations and translations. It often successfully segments the moving objects in the scene, even though such supervision is never provided.
References in corpus (2)
Cited by in corpus (90)
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- Image Matching across Wide Baselines: From Paper to Practice
- Monocular Depth Estimation Based On Deep Learning: An Overview
- Unsupervised Scale-consistent Depth Learning from Video
- Perception and Navigation in Autonomous Systems in the Era of Learning: A Survey
- DeepV2D: Video to Depth with Differentiable Structure from Motion
- Self-supervised Learning of Motion Capture
- GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
- Semantics for Robotic Mapping, Perception and Interaction: A Survey
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Movement science needs different pose tracking algorithms
- Unsupervised Video Object Segmentation for Deep Reinforcement Learning
- Neural Expectation Maximization
- MODNet: Moving Object Detection Network with Motion and Appearance for Autonomous Driving
- Video Object Segmentation and Tracking: A Survey
- Recurrent Neural Network for (Un-)supervised Learning of Monocular VideoVisual Odometry and Depth
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- BiFuse++: Self-supervised and Efficient Bi-projection Fusion for 360 Depth Estimation
- A Survey of Simultaneous Localization and Mapping with an Envision in 6G Wireless Networks
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- Every Pixel Counts ++: Joint Learning of Geometry and Motion with 3D Holistic Understanding
- Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera
- A Survey on Deep Learning Architectures for Image-based Depth Reconstruction
- Self-Supervised Deep Pose Corrections for Robust Visual Odometry
- Deep Reinforcement Learning in Computer Vision: A Comprehensive Survey
- LEGO: Learning Edge with Geometry all at Once by Watching Videos
- Learning 3D Dynamic Scene Representations for Robot Manipulation
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Occlusion Aware Unsupervised Learning of Optical Flow
- PeerNets: Exploiting Peer Wisdom Against Adversarial Attacks
- Unsupervised part representation by Flow Capsules
- EC-SfM: Efficient Covisibility-based Structure-from-Motion for Both Sequential and Unordered Images
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- Instance Embedding Transfer to Unsupervised Video Object Segmentation
- Robustness Meets Deep Learning: An End-to-End Hybrid Pipeline for Unsupervised Learning of Egomotion
- DFineNet: Ego-Motion Estimation and Depth Refinement from Sparse, Noisy Depth Input with RGB Guidance
- Sparse Representations for Object and Ego-motion Estimation in Dynamic Scenes
- Self-Supervised Learning of Depth and Camera Motion from 360° Videos
- Deep Global-Relative Networks for End-to-End 6-DoF Visual Localization and Odometry
- VisualEchoes: Spatial Image Representation Learning through Echolocation
- On the Coupling of Depth and Egomotion Networks for Self-Supervised Structure from Motion
- Every Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding
- The Edge of Depth: Explicit Constraints between Segmentation and Depth
- Large-Scale Object Discovery and Detector Adaptation from Unlabeled Video
- Self-Supervised Learning of Depth and Ego-motion with Differentiable Bundle Adjustment
- Consistent Video Depth Estimation
- InverseRenderNet: Learning single image inverse rendering
- Learning Independent Object Motion from Unlabelled Stereoscopic Videos
- Object category learning and retrieval with weak supervision
- Beyond Photometric Loss for Self-Supervised Ego-Motion Estimation
- Learning Correspondence from the Cycle-Consistency of Time
- Semantic-Guided Representation Enhancement for Self-supervised Monocular Trained Depth Estimation
- Unsupervised Learning of Depth and Ego-Motion from Cylindrical Panoramic Video with Applications for Virtual Reality
- TartanVO: A Generalizable Learning-based VO
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- DiPE: Deeper into Photometric Errors for Unsupervised Learning of Depth and Ego-motion from Monocular Videos
- Driven to Distraction: Self-Supervised Distractor Learning for Robust Monocular Visual Odometry in Urban Environments
- Layer-structured 3D Scene Inference via View Synthesis
- Instance-aware multi-object self-supervision for monocular depth prediction
- Multiview Supervision By Registration
- DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
- Toward Hierarchical Self-Supervised Monocular Absolute Depth Estimation for Autonomous Driving Applications
- On the Sins of Image Synthesis Loss for Self-supervised Depth Estimation
- Edge-Guided Occlusion Fading Reduction for a Light-Weighted Self-Supervised Monocular Depth Estimation
- Multiview Cross-supervision for Semantic Segmentation
- TransCamP: Graph Transformer for 6-DoF Camera Pose Estimation
- Unsupervised Event-based Learning of Optical Flow, Depth, and Egomotion
- Beyond the Camera: Neural Networks in World Coordinates
- Vision-based system identification and 3D keypoint discovery using dynamics constraints
- Robust Consistent Video Depth Estimation
- Learning by Inertia: Self-supervised Monocular Visual Odometry for Road Vehicles
- Self-Supervised Learning of Depth and Motion Under Photometric Inconsistency
- Sparse Coding Predicts Optic Flow Specificities of Zebrafish Pretectal Neurons
- Unsupervised Learning of Camera Pose with Compositional Re-estimation
- Neural Autonomous Navigation with Riemannian Motion Policy
- SIGNet: Semantic Instance Aided Unsupervised 3D Geometry Perception
- Cutting-edge 3D reconstruction solutions for underwater coral reef images: A review and comparison
- Depth Completion via Deep Basis Fitting
- Unsupervised Monocular Depth Perception: Focusing on Moving Objects
- IMU-Assisted Learning of Single-View Rolling Shutter Correction
- Radar-Camera Pixel Depth Association for Depth Completion
- Deep Coarse-to-fine Dense Light Field Reconstruction with Flexible Sampling and Geometry-aware Fusion
- Learning Single-Image Depth from Videos using Quality Assessment Networks
- PLG-IN: Pluggable Geometric Consistency Loss with Wasserstein Distance in Monocular Depth Estimation
- Outdoor inverse rendering from a single image using multiview self-supervision
- RGB-based Semantic Segmentation Using Self-Supervised Depth Pre-Training
- 3D Surface Reconstruction by Pointillism
- Indoor GeoNet: Weakly Supervised Hybrid Learning for Depth and Pose Estimation
- Unpaired Single-Image Depth Synthesis with cycle-consistent Wasserstein GANs
- Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space