Unsupervised Learning of Video Representations using LSTMs
arXiv:1502.04681
Abstract
We use multilayer Long Short Term Memory (LSTM) networks to learn representations of video sequences. Our model uses an encoder LSTM to map an input sequence into a fixed length representation. This representation is decoded using single or multiple decoder LSTMs to perform different tasks, such as reconstructing the input sequence, or predicting the future sequence. We experiment with two kinds of input sequences - patches of image pixels and high-level representations ("percepts") of video frames extracted using a pretrained convolutional net. We explore different design choices such as whether the decoder LSTMs should condition on the generated output. We analyze the outputs of the model qualitatively to see how well the model can extrapolate the learned video representation into the future and into the past. We try to visualize and interpret the learned features. We stress test the model by running it on longer time scales and on out-of-domain data. We further evaluate the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets. We show that the representations help improve classification accuracy, especially when there are only a few training examples. Even models pretrained on unrelated datasets (300 hours of YouTube videos) can help action recognition performance.
Added link to code on github
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Recurrent Neural Network Regularization
- DRAW: A Recurrent Neural Network For Image Generation
- Video (language) modeling: a baseline for generative models of natural videos
Cited by in corpus (313)
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- A Comprehensive Survey on Graph Anomaly Detection with Deep Learning
- What Makes for Good Views for Contrastive Learning?
- Human Action Recognition from Various Data Modalities: A Review
- Semi-supervised Sequence Learning
- Image and Video Compression with Neural Networks: A Review
- Exploiting Feature and Class Relationships in Video Categorization with Regularized Deep Neural Networks
- Action Recognition using Visual Attention
- Deep and Confident Prediction for Time Series at Uber
- Fast and Furious: Real Time End-to-End 3D Detection, Tracking and Motion Forecasting with a Single Convolutional Net
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Neural Encoding and Decoding with Deep Learning for Dynamic Natural Vision
- Rank Pooling for Action Recognition
- A Review on Deep Learning Techniques for Video Prediction
- Hard Negative Mixing for Contrastive Learning
- Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Spatio-temporal video autoencoder with differentiable memory
- Unsupervised Scalable Representation Learning for Multivariate Time Series
- Motion-Attentive Transition for Zero-Shot Video Object Segmentation
- Stochastic Adversarial Video Prediction
- Overcoming Limitations of Mixture Density Networks: A Sampling and Fitting Framework for Multimodal Future Prediction
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- TSM: Temporal Shift Module for Efficient Video Understanding
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model
- Adversarial Video Generation on Complex Datasets
- Recurrent Neural Networks: An Embedded Computing Perspective
- Temporal Localization of Fine-Grained Actions in Videos by Domain Transfer from Web Images
- 3D PersonVLAD: Learning Deep Global Representations for Video-based Person Re-identification
- Recurrent Mixture Density Network for Spatiotemporal Visual Attention
- Land Cover Classification from Multi-temporal, Multi-spectral Remotely Sensed Imagery using Patch-Based Recurrent Neural Networks
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Recurrent Environment Simulators
- Augmenting Physical Models with Deep Networks for Complex Dynamics Forecasting
- Going in circles is the way forward: the role of recurrence in visual inference
- A critical analysis of self-supervision, or what we can learn from a single image
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Few-Shot Deep Adversarial Learning for Video-based Person Re-identification
- Generative Image Modeling Using Spatial LSTMs
- Self-labelling via simultaneous clustering and representation learning
- Graph Neural Network and Spatiotemporal Transformer Attention for 3D Video Object Detection from Point Clouds
- Investigating Pose Representations and Motion Contexts Modeling for 3D Motion Prediction
- 2022 Review of Data-Driven Plasma Science
- Convolutional Tensor-Train LSTM for Spatio-temporal Learning
- Spatiotemporal Contrastive Video Representation Learning
- PointRNN: Point Recurrent Neural Network for Moving Point Cloud Processing
- Analyzing Human-Human Interactions: A Survey
- Long-term Forecasting using Higher Order Tensor RNNs
- Convolutional Invasion and Expansion Networks for Tumor Growth Prediction
- Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching
- Spatio-Temporal Neural Networks for Space-Time Series Forecasting and Relations Discovery
- VideoLSTM Convolves, Attends and Flows for Action Recognition
- Action-Attending Graphic Neural Network
- Symbolic Pregression: Discovering Physical Laws from Distorted Video
- Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification
- Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Temporal Modeling Approaches for Large-scale Youtube-8M Video Understanding
- Dual Convolutional LSTM Network for Referring Image Segmentation
- Video Representation Learning by Dense Predictive Coding
- Uncovering Temporal Context for Video Question and Answering
- Generative Models for Low-Rank Video Representation and Reconstruction
- Attention-based Fully Gated CNN-BGRU for Russian Handwritten Text
- Combining Static and Dynamic Features for Multivariate Sequence Classification
- Predictive-Corrective Networks for Action Detection
- Semantic Image Networks for Human Action Recognition
- Watching the World Go By: Representation Learning from Unlabeled Videos
- Semi-supervised Federated Learning for Activity Recognition
- Temporally smooth online action detection using cycle-consistent future anticipation
- PHD-GIFs: Personalized Highlight Detection for Automatic GIF Creation
- Visual Dynamics: Stochastic Future Generation via Layered Cross Convolutional Networks
- Comparing recurrent and convolutional neural networks for predicting wave propagation
- RGB-D-based Human Motion Recognition with Deep Learning: A Survey
- SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving
- Generative Adversarial Networks for Image and Video Synthesis: Algorithms and Applications
- Forecasting future action sequences with attention: a new approach to weakly supervised action forecasting
- Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses
- Classification of Handwritten Names of Cities and Handwritten Text Recognition using Various Deep Learning Models
- Predicting Visual Context for Unsupervised Event Segmentation in Continuous Photo-streams
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
- Short-term daily precipitation forecasting with seasonally-integrated autoencoder
- Self-supervised Learning for Large-scale Item Recommendations
- FitVid: Overfitting in Pixel-Level Video Prediction
- HybridNet: Integrating Model-based and Data-driven Learning to Predict Evolution of Dynamical Systems
- Continual Learning with Gated Incremental Memories for sequential data processing
- Unsupervised Deep Anomaly Detection for Multi-Sensor Time-Series Signals
- Video-based Human Action Recognition using Deep Learning: A Review
- Photo-Realistic Video Prediction on Natural Videos of Largely Changing Frames
- KOHTD: Kazakh Offline Handwritten Text Dataset
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Memory In Memory: A Predictive Neural Network for Learning Higher-Order Non-Stationarity from Spatiotemporal Dynamics
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative Models
- STWalk: Learning Trajectory Representations in Temporal Graphs
- Transitive Invariance for Self-supervised Visual Representation Learning
- CloudCast: A Satellite-Based Dataset and Baseline for Forecasting Clouds
- Pointly-Supervised Action Localization
- MAD: Self-Supervised Masked Anomaly Detection Task for Multivariate Time Series
- Deep Learning for Vision-based Prediction: A Survey
- From Here to There: Video Inbetweening Using Direct 3D Convolutions
- Siamese Networks for Weakly Supervised Human Activity Recognition
- High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
- Self-supervised Human Activity Recognition by Learning to Predict Cross-Dimensional Motion
- 3D Human Action Representation Learning via Cross-View Consistency Pursuit
- Depth2Action: Exploring Embedded Depth for Large-Scale Action Recognition
- A neural network trained to predict future video frames mimics critical properties of biological neuronal responses and perception
- Gait Recognition via Disentangled Representation Learning
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Precipitation Nowcasting with Star-Bridge Networks
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- PredNet and Predictive Coding: A Critical Review
- Unsupervised Learning of Object Structure and Dynamics from Videos
- Spatiotemporal Recurrent Convolutional Networks for Recognizing Spontaneous Micro-expressions
- High-Dimensional Similarity Search with Quantum-Assisted Variational Autoencoder
- Depth Pooling Based Large-scale 3D Action Recognition with Convolutional Neural Networks
- Recomposition vs. Prediction: A Novel Anomaly Detection for Discrete Events Based On Autoencoder
- Unsupervised Learning of Video Representations via Dense Trajectory Clustering
- Skeleton Cloud Colorization for Unsupervised 3D Action Representation Learning
- Consistent Generative Query Networks
- Scaling Autoregressive Video Models
- Mobile Video Action Recognition
- Memory Warps for Learning Long-Term Online Video Representations
- A Review of Artificial Intelligence Technologies for Early Prediction of Alzheimer's Disease
- On Learning Disentangled Representations for Gait Recognition
- Self-supervised Point Cloud Prediction Using 3D Spatio-temporal Convolutional Networks
- Video Generation from Single Semantic Label Map
- Deep Analysis of CNN-based Spatio-temporal Representations for Action Recognition
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Learning audio sequence representations for acoustic event classification
- Influence-guided Data Augmentation for Neural Tensor Completion
- Self-supervised Modal and View Invariant Feature Learning
- Reduced-Gate Convolutional LSTM Using Predictive Coding for Spatiotemporal Prediction
- Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- Clockwork Variational Autoencoders
- Gated networks: an inventory
- Effect of Architectures and Training Methods on the Performance of Learned Video Frame Prediction
- Transformers predicting the future. Applying attention in next-frame and time series forecasting
- Learning Regularity in Skeleton Trajectories for Anomaly Detection in Videos
- Mutual Suppression Network for Video Prediction using Disentangled Features
- PDE-Driven Spatiotemporal Disentanglement
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Unsupervised Motion Representation Learning with Capsule Autoencoders
- Pose Guided Human Video Generation
- A Time Attention based Fraud Transaction Detection Framework
- Spatio-temporal Manifold Learning for Human Motions via Long-horizon Modeling
- Skeleton-based Gait Index Estimation with LSTMs
- Interactive Fusion of Multi-level Features for Compositional Activity Recognition
- FASTER Recurrent Networks for Efficient Video Classification
- MotionRNN: A Flexible Model for Video Prediction with Spacetime-Varying Motions
- Sub-Seasonal Climate Forecasting via Machine Learning: Challenges, Analysis, and Advances
- DwNet: Dense warp-based network for pose-guided human video generation
- Towards a Near Universal Time Series Data Mining Tool: Introducing the Matrix Profile
- Compressed Video Action Recognition with Refined Motion Vector
- Revisiting Hierarchical Approach for Persistent Long-Term Video Prediction
- Object-Centric Representation Learning from Unlabeled Videos
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Multi-Scale Video Frame-Synthesis Network with Transitive Consistency Loss
- Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient Data
- Learning Representations for Predicting Future Activities
- Diverse Image Synthesis from Semantic Layouts via Conditional IMLE
- Learning representations for multivariate time series with missing data using Temporal Kernelized Autoencoders
- Order Matters: Shuffling Sequence Generation for Video Prediction
- An Encoder-Decoder Based Approach for Anomaly Detection with Application in Additive Manufacturing
- DROCC: Deep Robust One-Class Classification
- Video Contents Understanding using Deep Neural Networks
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Planning Robot Motion using Deep Visual Prediction
- Group-Skeleton-Based Human Action Recognition in Complex Events
- MVC-Net: A Convolutional Neural Network Architecture for Manifold-Valued Images With Applications
- TwoStreamVAN: Improving Motion Modeling in Video Generation
- Unsupervised Bi-directional Flow-based Video Generation from one Snapshot
- Deep Multimodal Feature Encoding for Video Ordering
- Adversarial Self-Supervised Learning for Semi-Supervised 3D Action Recognition
- Self-Supervised Multi-View Learning via Auto-Encoding 3D Transformations
- Simple Video Generation using Neural ODEs
- Surprisal-Triggered Conditional Computation with Neural Networks
- BézierSketch: A generative model for scalable vector sketches
- Improving Sequential Latent Variable Models with Autoregressive Flows
- Learning without Prejudice: Avoiding Bias in Webly-Supervised Action Recognition
- Precipitation nowcasting using a stochastic variational frame predictor with learned prior distribution
- Evolving Losses for Unlabeled Video Representation Learning
- Attentioned Convolutional LSTM InpaintingNetwork for Anomaly Detection in Videos
- Unsupervised Representation Learning of DNA Sequences
- Temporal View Synthesis of Dynamic Scenes through 3D Object Motion Estimation with Multi-Plane Images
- What Would You Do? Acting by Learning to Predict
- Learning Temporal Transformations From Time-Lapse Videos
- Multi-View Frame Reconstruction with Conditional GAN
- Modeling Spatio-Temporal Human Track Structure for Action Localization
- FDNet: A Deep Learning Approach with Two Parallel Cross Encoding Pathways for Precipitation Nowcasting
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- Target-Embedding Autoencoders for Supervised Representation Learning
- Self-Paced Video Data Augmentation with Dynamic Images Generated by Generative Adversarial Networks
- Review of Video Predictive Understanding: Early Action Recognition and Future Action Prediction
- Temporal Cross-Media Retrieval with Soft-Smoothing
- Learning from Videos with Deep Convolutional LSTM Networks
- Frame-To-Frame Consistent Semantic Segmentation
- Future Segmentation Using 3D Structure
- Driver Intention Anticipation Based on In-Cabin and Driving Scene Monitoring
- Future Frame Prediction of a Video Sequence
- Variational Tracking and Prediction with Generative Disentangled State-Space Models
- A Proposal-based Approach for Activity Image-to-Video Retrieval
- Learning Monocular Visual Odometry via Self-Supervised Long-Term Modeling
- PREDICT & CLUSTER: Unsupervised Skeleton Based Action Recognition
- World-Consistent Video-to-Video Synthesis
- One-Shot Imitation Filming of Human Motion Videos
- Video Reenactment as Inductive Bias for Content-Motion Disentanglement
- Dynamic Filtering with Large Sampling Field for ConvNets
- Self-Supervised Multi-Object Tracking with Cross-Input Consistency
- Learning Cross-modal Contrastive Features for Video Domain Adaptation
- MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish
- MoNet: Motion-based Point Cloud Prediction Network
- Spatio-Temporal LSTM with Trust Gates for 3D Human Action Recognition
- Dilated Convolutional Neural Networks for Sequential Manifold-valued Data
- Customizing Sequence Generation with Multi-Task Dynamical Systems
- Understanding the Perceived Quality of Video Predictions
- Learning Spatio-Temporal Representation with Local and Global Diffusion
- ManifoldNorm: Extending normalizations on Riemannian Manifolds
- Future Frame Prediction for Robot-assisted Surgery
- 3D Human motion anticipation and classification
- Wide and Narrow: Video Prediction from Context and Motion
- Diverse Video Generation using a Gaussian Process Trigger
- iMiGUE: An Identity-free Video Dataset for Micro-Gesture Understanding and Emotion Analysis
- Unsupervised detection of mouse behavioural anomalies using two-stream convolutional autoencoders
- OmniPrint: A Configurable Printed Character Synthesizer
- Contrast-reconstruction Representation Learning for Self-supervised Skeleton-based Action Recognition
- Long-Short Temporal Contrastive Learning of Video Transformers
- Learning Temporal Dynamics from Cycles in Narrated Video
- From Recognition to Prediction: Analysis of Human Action and Trajectory Prediction in Video
- Latent Neural Differential Equations for Video Generation
- Learning Disentangled Representations of Video with Missing Data
- Amortized Population Gibbs Samplers with Neural Sufficient Statistics
- On the difficulty of learning and predicting the long-term dynamics of bouncing objects
- Object Localization with a Weakly Supervised CapsNet
- Particle Filter Bridge Interpolation
- Hierarchical Video Generation for Complex Data
- Learning Dynamical Systems from Noisy Sensor Measurements using Multiple Shooting
- Future Video Synthesis with Object Motion Prediction
- Sampling-free Uncertainty Estimation in Gated Recurrent Units with Exponential Families
- A Neural Architecture for Detecting Confusion in Eye-tracking Data
- Predicting the Future with Transformational States
- Self-Supervised Video Representation Learning by Video Incoherence Detection
- Bridging the Gap Between Training and Inference for Spatio-Temporal Forecasting
- Self-Supervised Decomposition, Disentanglement and Prediction of Video Sequences while Interpreting Dynamics: A Koopman Perspective
- Contrastive Video Representation Learning via Adversarial Perturbations
- Sound2Sight: Generating Visual Dynamics from Sound and Context
- DeepLandscape: Adversarial Modeling of Landscape Video
- Causal Future Prediction in a Minkowski Space-Time
- ModeRNN: Harnessing Spatiotemporal Mode Collapse in Unsupervised Predictive Learning
- Encode the Unseen: Predictive Video Hashing for Scalable Mid-Stream Retrieval
- Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition
- Interpretable Intuitive Physics Model
- Learning Representations from Deep Networks Using Mode Synthesizers
- Towards Good Practices of U-Net for Traffic Forecasting
- Learning Scene Dynamics from Point Cloud Sequences
- Early Prediction of Alzheimer's Disease Dementia Based on Baseline Hippocampal MRI and 1-Year Follow-Up Cognitive Measures Using Deep Recurrent Neural Networks
- Disentangling Video with Independent Prediction
- Self-supervised Learning with Fully Convolutional Networks
- Video Summarization via Actionness Ranking
- Estimating Emotional Intensity from Body Poses for Human-Robot Interaction
- LMVP: Video Predictor with Leaked Motion Information
- Learning to Align Sequential Actions in the Wild
- Video Contrastive Learning with Global Context
- Questions to Guide the Future of Artificial Intelligence Research
- Autonomous Learning of Features for Control: Experiments with Embodied and Situated Agents
- Multi-Decoder RNN Autoencoder Based on Variational Bayes Method
- Dynamic Variational Autoencoders for Visual Process Modeling
- Lyric Video Analysis Using Text Detection and Tracking
- Understanding Road Layout from Videos as a Whole
- STEP-GAN: A Step-by-Step Training for Multi Generator GANs with application to Cyber Security in Power Systems
- Brain-Inspired Inference on Missing Video Sequence
- MoEVC: A Mixture-of-experts Voice Conversion System with Sparse Gating Mechanism for Accelerating Online Computation
- Machine Learning based Post Processing Artifact Reduction in HEVC Intra Coding
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- Learning to navigate image manifolds induced by generative adversarial networks for unsupervised video generation
- Cubic LSTMs for Video Prediction
- Implicit Label Augmentation on Partially Annotated Clips via Temporally-Adaptive Features Learning
- Exploring Temporal Information for Improved Video Understanding
- Predicting Future Opioid Incidences Today
- Deep Multi-Kernel Convolutional LSTM Networks and an Attention-Based Mechanism for Videos
- Affine-modeled video extraction from a single motion blurred image
- Temporally Consistent Depth Prediction with Flow-Guided Memory Units
- Action Anticipation with RBF Kernelized Feature Mapping RNN
- Generating Videos of Zero-Shot Compositions of Actions and Objects
- Visual Reaction: Learning to Play Catch with Your Drone
- Point-to-Point Video Generation
- A Multi-view Perspective of Self-supervised Learning
- Learning to infer in recurrent biological networks
- Deep Sequence Learning for Video Anticipation: From Discrete and Deterministic to Continuous and Stochastic
- PARIS: Personalized Activity Recommendation for Improving Sleep Quality
- SDCNet: Video Prediction Using Spatially-Displaced Convolution
- Taylor saves for later: disentanglement for video prediction using Taylor representation
- Vision-Guided Forecasting -- Visual Context for Multi-Horizon Time Series Forecasting
- Towards an Interpretable Latent Space in Structured Models for Video Prediction
- Deep Photovoltaic Nowcasting
- Multi-Modal Temporal Convolutional Network for Anticipating Actions in Egocentric Videos
- From Single to Multiple: Leveraging Multi-level Prediction Spaces for Video Forecasting
- Stochastic Dynamics for Video Infilling
- Speech Representations and Phoneme Classification for Preserving the Endangered Language of Ladin
- Incorporating Scalability in Unsupervised Spatio-Temporal Feature Learning
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- A Framework for Multisensory Foresight for Embodied Agents
- Stochastic Talking Face Generation Using Latent Distribution Matching
- Video SemNet: Memory-Augmented Video Semantic Network
- Smoothed Gaussian Mixture Models for Video Classification and Recommendation
- Knowledge as Invariance -- History and Perspectives of Knowledge-augmented Machine Learning
- Probability Trajectory: One New Movement Description for Trajectory Prediction
- Understanding in Artificial Intelligence
- CLTA: Contents and Length-based Temporal Attention for Few-shot Action Recognition
- Discriminative Video Representation Learning Using Support Vector Classifiers
- Unsupervised learning of the brain connectivity dynamic using residual D-net