Describing Videos by Exploiting Temporal Structure
arXiv:1502.08029
Abstract
Recent progress in using recurrent neural networks (RNNs) for image description has motivated the exploration of their application for video description. However, while images are static, working with videos requires modeling their dynamic temporal structure and then properly integrating that information into a natural language description. In this context, we propose an approach that successfully takes into account both the local and global temporal structure of videos to produce descriptions. First, our approach incorporates a spatial temporal 3-D convolutional neural network (3-D CNN) representation of the short temporal dynamics. The 3-D CNN representation is trained on video action recognition tasks, so as to produce a representation that is tuned to human motion and behavior. Second we propose a temporal attention mechanism that allows to go beyond local temporal modeling and learns to automatically select the most relevant temporal segments given the text-generating RNN. Our approach exceeds the current state-of-art for both BLEU and METEOR metrics on the Youtube2Text dataset. We also present results on a new, larger and more challenging dataset of paired video and natural language descriptions.
Accepted to ICCV15. This version comes with code release and supplementary material
References in corpus (15)
- Deep Learning in Neural Networks: An Overview
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- ADADELTA: An Adaptive Learning Rate Method
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Recurrent Neural Network Regularization
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Theano: new features and speed improvements
- Video (language) modeling: a baseline for generative models of natural videos
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- Deep Visual-Semantic Alignments for Generating Image Descriptions
- Weakly Supervised Action Labeling in Videos Under Ordering Constraints
Cited by in corpus (110)
- Recent Advances in Recurrent Neural Networks
- Action Recognition using Visual Attention
- ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks
- Learning to diagnose from scratch by exploiting dependencies among labels
- Sequence to Sequence -- Video to Text
- Delving Deeper into Convolutional Networks for Learning Video Representations
- Flow-Guided Feature Aggregation for Video Object Detection
- Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Story Generation from Sequence of Independent Short Descriptions
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- Recurrent Batch Normalization
- First Step toward Model-Free, Anonymous Object Tracking with Recurrent Neural Networks
- ExtremeWeather: A large-scale climate dataset for semi-supervised detection, localization, and understanding of extreme weather events
- Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
- Video-based Sign Language Recognition without Temporal Segmentation
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- Visual Question Answering: A Survey of Methods and Datasets
- DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks
- End-to-End Dense Video Captioning with Masked Transformer
- Reconstruction Network for Video Captioning
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Jointly Localizing and Describing Events for Dense Video Captioning
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- What value do explicit high level concepts have in vision to language problems?
- A Recurrent Encoder-Decoder Network for Sequential Face Alignment
- Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Video Summarization using Deep Semantic Features
- Weakly Supervised Dense Video Captioning
- Less Is More: Picking Informative Frames for Video Captioning
- V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Attention-Based Multimodal Fusion for Video Description
- Video Captioning via Hierarchical Reinforcement Learning
- Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
- Multi-granularity Generator for Temporal Action Proposal
- Memory-augmented Attention Modelling for Videos
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Generating Descriptions with Grounded and Co-Referenced People
- From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning
- Video Captioning with Transferred Semantic Attributes
- Bidirectional Long-Short Term Memory for Video Description
- Video Fill in the Blank with Merging LSTMs
- Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning
- Memory Warps for Learning Long-Term Online Video Representations
- End-to-end Video-level Representation Learning for Action Recognition
- Action Recognition with Joint Attention on Multi-Level Deep Features
- A Hierarchical Approach for Generating Descriptive Image Paragraphs
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- Integrating both Visual and Audio Cues for Enhanced Video Caption
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- Collaborative Summarization of Topic-Related Videos
- FASTER Recurrent Networks for Efficient Video Classification
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention
- Context, Attention and Audio Feature Explorations for Audio Visual Scene-Aware Dialog
- Title Generation for User Generated Videos
- Attacking Automatic Video Analysis Algorithms: A Case Study of Google Cloud Video Intelligence API
- A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
- Recurrent Attention Models for Depth-Based Person Identification
- Multimodal Memory Modelling for Video Captioning
- Move Forward and Tell: A Progressive Generator of Video Descriptions
- Video Captioning with Guidance of Multimodal Latent Topics
- A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
- Top-down Visual Saliency Guided by Captions
- Large-scale Video Classification guided by Batch Normalized LSTM Translator
- Modeling Multimodal Clues in a Hybrid Deep Learning Framework for Video Classification
- Attention-GAN for Object Transfiguration in Wild Images
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Action Classification and Highlighting in Videos
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- An Attention-Based Approach for Single Image Super Resolution
- Grounded Objects and Interactions for Video Captioning
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Early Improving Recurrent Elastic Highway Network
- Temporal Tessellation: A Unified Approach for Video Analysis
- RUC+CMU: System Report for Dense Captioning Events in Videos
- Fast Object Localization Using a CNN Feature Map Based Multi-Scale Search
- MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Federated Multi-task Hierarchical Attention Model for Sensor Analytics
- Action Tubelet Detector for Spatio-Temporal Action Localization
- Mining for meaning: from vision to language through multiple networks consensus
- Temporal Attention-Gated Model for Robust Sequence Classification
- Grounded Video Description
- X-modaler: A Versatile and High-performance Codebase for Cross-modal Analytics
- Context-aware Cascade Attention-based RNN for Video Emotion Recognition
- Improving Classification by Improving Labelling: Introducing Probabilistic Multi-Label Object Interaction Recognition
- Anticipating Daily Intention using On-Wrist Motion Triggered Sensing
- AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos
- Learning Joint Representations of Videos and Sentences with Web Image Search
- Generating Video Descriptions with Topic Guidance
- Learning Articulated Motion Models from Visual and Lingual Signals
- Reinforced Temporal Attention and Split-Rate Transfer for Depth-Based Person Re-Identification
- Middle-Out Decoding
- Multi-modal Dense Video Captioning
- Towards Diverse Paragraph Captioning for Untrimmed Videos
- Multi-View Surveillance Video Summarization via Joint Embedding and Sparse Optimization
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Knowledge Fusion Transformers for Video Action Recognition
- Curvature: A signature for Action Recognition in Video Sequences
- The Long-Short Story of Movie Description