Spatio-temporal video autoencoder with differentiable memory
arXiv:1511.06309
Abstract
We describe a new spatio-temporal video autoencoder, based on a classic spatial image autoencoder and a novel nested temporal autoencoder. The temporal encoder is represented by a differentiable visual memory composed of convolutional long short-term memory (LSTM) cells that integrate changes over time. Here we target motion changes and use as temporal decoder a robust optical flow prediction module together with an image sampler serving as built-in feedback loop. The architecture is end-to-end differentiable. At each time step, the system receives as input a video frame, predicts the optical flow based on the current observation and the LSTM memory state as a dense transformation map, and applies it to the current frame to generate the next frame. By minimising the reconstruction error between the predicted next frame and the corresponding ground truth next frame, we train the whole system to extract features useful for motion estimation without any supervision effort. We present one direct application of the proposed framework in weakly-supervised semantic segmentation of videos through label propagation using optical flow.
The experiments section has been extended and a direct application to weakly-supervised video segmentation through label propagation has been included
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- Unsupervised Learning of Video Representations using LSTMs
- Long Short-Term Memory Based Recurrent Neural Network Architectures for Large Vocabulary Speech Recognition
- Learning to See by Moving
- SceneNet: Understanding Real World Indoor Scenes With Synthetic Data
- rnn : Recurrent Library for Torch
Cited by in corpus (65)
- Dynamic Filter Networks
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- SfM-Net: Learning of Structure and Motion from Video
- A Review on Deep Learning Techniques for Video Prediction
- Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
- Combining Fully Convolutional and Recurrent Neural Networks for 3D Biomedical Image Segmentation
- Anomaly Detection in Video Using Predictive Convolutional Long Short-Term Memory Networks
- Self-supervised Learning of Motion Capture
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Recurrent Environment Simulators
- Self-supervised learning of a facial attribute embedding from video
- Learning to Decompose and Disentangle Representations for Video Prediction
- Unsupervised Learning of View-invariant Action Representations
- Transformation-Based Models of Video Sequences
- Learning Video Object Segmentation with Visual Memory
- Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation
- Mobile Video Object Detection with Temporally-Aware Feature Maps
- Predicting Deeper into the Future of Semantic Segmentation
- Spatio-temporal Stacked LSTM for Temperature Prediction in Weather Forecasting
- Unsupervised Monocular Depth Estimation with Left-Right Consistency
- Back to Basics: Unsupervised Learning of Optical Flow via Brightness Constancy and Motion Smoothness
- End-to-end Flow Correlation Tracking with Spatial-temporal Attention
- Unsupervised Deep Anomaly Detection for Multi-Sensor Time-Series Signals
- Predicting Video with VQVAE
- Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics
- Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
- Occlusion Aware Unsupervised Learning of Optical Flow
- Video Ladder Networks
- Inception-inspired LSTM for Next-frame Video Prediction
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- PSIque: Next Sequence Prediction of Satellite Images using a Convolutional Sequence-to-Sequence Network
- Scaling Autoregressive Video Models
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Deep Tracking on the Move: Learning to Track the World from a Moving Vehicle using Recurrent Neural Networks
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Motion Prediction Under Multimodality with Conditional Stochastic Networks
- CortexNet: a Generic Network Family for Robust Visual Temporal Representations
- Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation
- Temporal Interpolation via Motion Field Prediction
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient Data
- Anomaly Detection using Deep Reconstruction and Forecasting for Autonomous Systems
- Order Matters: Shuffling Sequence Generation for Video Prediction
- End-to-End Learning for Image Burst Deblurring
- Adversarial Video Compression Guided by Soft Edge Detection
- Unsupervised Bi-directional Flow-based Video Generation from one Snapshot
- A Neurally-Inspired Hierarchical Prediction Network for Spatiotemporal Sequence Learning and Prediction
- DriveGuard: Robustification of Automated Driving Systems with Deep Spatio-Temporal Convolutional Autoencoder
- What Would You Do? Acting by Learning to Predict
- Better Guider Predicts Future Better: Difference Guided Generative Adversarial Networks
- Long-Term Image Boundary Prediction
- Object Localization with a Weakly Supervised CapsNet
- Wide and Narrow: Video Prediction from Context and Motion
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- Every Frame Counts: Joint Learning of Video Segmentation and Optical Flow
- Unsupervised Learning-based Depth Estimation aided Visual SLAM Approach
- LMVP: Video Predictor with Leaked Motion Information
- Future Video Synthesis with Object Motion Prediction
- Recurrent Flow-Guided Semantic Forecasting
- Occlusion Aware Unsupervised Learning of Optical Flow From Video
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Affine-modeled video extraction from a single motion blurred image
- A Variational Auto-Encoder Model for Stochastic Point Processes
- Irregular Convolutional Auto-Encoder on Point Clouds