Delving Deeper into Convolutional Networks for Learning Video Representations
arXiv:1511.06432
Abstract
We propose an approach to learn spatio-temporal features in videos from intermediate visual representations we call "percepts" using Gated-Recurrent-Unit Recurrent Networks (GRUs).Our method relies on percepts that are extracted from all level of a deep convolutional network trained on the large ImageNet dataset. While high-level percepts contain highly discriminative information, they tend to have a low-spatial resolution. Low-level percepts, on the other hand, preserve a higher spatial resolution from which we can model finer motion patterns. Using low-level percepts can leads to high-dimensionality video representations. To mitigate this effect and control the model number of parameters, we introduce a variant of the GRU model that leverages the convolution operations to enforce sparse connectivity of the model units and share parameters across the input spatial locations. We empirically validate our approach on both Human Action Recognition and Video Captioning tasks. In particular, we achieve results equivalent to state-of-art on the YouTube2Text dataset using a simpler text-decoder model and without extra 3D CNN features.
ICLR 2016
References in corpus (18)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting
- ADADELTA: An Adaptive Learning Rate Method
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Going Deeper with Convolutions
- Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors
- Theano: new features and speed improvements
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Beyond Short Snippets: Deep Networks for Video Classification
- Describing Videos by Exploiting Temporal Structure
- Learning Spatiotemporal Features with 3D Convolutional Networks
- CIDEr: Consensus-based Image Description Evaluation
- Hierarchical Recurrent Neural Encoder for Video Representation with Application to Captioning
- Beyond Gaussian Pyramid: Multi-skip Feature Stacking for Action Recognition
Cited by in corpus (46)
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- Faster Mean-shift: GPU-accelerated clustering for cosine embedding-based cell segmentation and tracking
- Adversarial Video Generation on Complex Datasets
- Flow-Guided Feature Aggregation for Video Object Detection
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- A Hybrid Spatial-temporal Deep Learning Architecture for Lane Detection
- Learning Video Object Segmentation with Visual Memory
- Reconstruction Network for Video Captioning
- Towards High Performance Video Object Detection for Mobiles
- Seamless lightning nowcasting with recurrent-convolutional deep learning
- Impression Network for Video Object Detection
- Spatio-Temporal Attention Models for Grounded Video Captioning
- 3D Gated Recurrent Fusion for Semantic Scene Completion
- Semantic Compositional Networks for Visual Captioning
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Transformation-based Adversarial Video Prediction on Large-Scale Data
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- Convolutional Residual Memory Networks
- A Time-domain Monaural Speech Enhancement with Feedback Learning
- Multimodal Memory Modelling for Video Captioning
- Object Detection in Videos by High Quality Object Linking
- Object Discovery in Videos as Foreground Motion Clustering
- A Recursive Network with Dynamic Attention for Monaural Speech Enhancement
- Self-supervised Video Representation Learning by Context and Motion Decoupling
- A Fusion Approach for Multi-Frame Optical Flow Estimation
- 3D Robot Pose Estimation from 2D Images
- Beyond One Glance: Gated Recurrent Architecture for Hand Segmentation
- Enhancing Salient Object Segmentation Through Attention
- Cube Padding for Weakly-Supervised Saliency Prediction in 360° Videos
- Deep Optical Flow Estimation Via Multi-Scale Correspondence Structure Learning
- PoseConvGRU: A Monocular Approach for Visual Ego-motion Estimation by Learning
- Sequential anatomy localization in fetal echocardiography videos
- Brain-Inspired Deep Imitation Learning for Autonomous Driving Systems
- Hierarchical Video Generation for Complex Data
- Segmenting Medical MRI via Recurrent Decoding Cell
- Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning
- Efficient Spatialtemporal Context Modeling for Action Recognition
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Tracking the Untrackable
- LiDAR-based Recurrent 3D Semantic Segmentation with Temporal Memory Alignment
- Automatic ultrasound vessel segmentation with deep spatiotemporal context learning
- Cine-MRI detection of abdominal adhesions with spatio-temporal deep learning
- Real-time Linear Operator Construction and State Estimation with the Kalman Filter
- Constrained-size Tensorflow Models for YouTube-8M Video Understanding Challenge
- Object Parsing in Sequences Using CoordConv Gated Recurrent Networks