Sequence to Sequence -- Video to Text
arXiv:1505.00487
Abstract
Real-world videos often have complex dynamics; and methods for generating open-domain video descriptions should be sensitive to temporal structure and allow both input (sequence of frames) and output (sequence of words) of variable length. To approach this problem, we propose a novel end-to-end sequence-to-sequence model to generate captions for videos. For this we exploit recurrent neural networks, specifically LSTMs, which have demonstrated state-of-the-art performance in image caption generation. Our LSTM model is trained on video-sentence pairs and learns to associate a sequence of video frames to a sequence of words in order to generate a description of the event in the video clip. Our model naturally is able to learn the temporal structure of the sequence of frames as well as the sequence model of the generated sentences, i.e. a language model. We evaluate several variants of our model that exploit different visual features on a standard set of YouTube videos and two movie description datasets (M-VAD and MPII-MD).
ICCV 2015 camera-ready. Includes code, project page and LSMDC challenge results
References in corpus (15)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Going Deeper with Convolutions
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- Learning to Execute
- Describing Videos by Exploiting Temporal Structure
- Learning a Recurrent Visual Representation for Image Caption Generation
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- A Dataset for Movie Description
- Finding Action Tubes
Cited by in corpus (104)
- A Critical Review of Recurrent Neural Networks for Sequence Learning
- Network Trimming: A Data-Driven Neuron Pruning Approach towards Efficient Deep Architectures
- Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond
- Convolutional Two-Stream Network Fusion for Video Action Recognition
- Self-Supervised Video Hashing with Hierarchical Binary Auto-encoder
- Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
- Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
- Frame- and Segment-Level Features and Candidate Pool Evaluation for Video Caption Generation
- Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Video Summarization with Long Short-term Memory
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- Video-based Sign Language Recognition without Temporal Segmentation
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- End-to-End Dense Video Captioning with Masked Transformer
- Reconstruction Network for Video Captioning
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning
- Jointly Localizing and Describing Events for Dense Video Captioning
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
- Forecasting future action sequences with attention: a new approach to weakly supervised action forecasting
- Show, Adapt and Tell: Adversarial Training of Cross-domain Image Captioner
- Jointly Modeling Embedding and Translation to Bridge Video and Language
- Weakly Supervised Dense Video Captioning
- Connectionist Temporal Modeling for Weakly Supervised Action Labeling
- Spatio-Temporal Attention Models for Grounded Video Captioning
- Video Captioning via Hierarchical Reinforcement Learning
- Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
- Learning Word Embeddings from Speech
- Memory-augmented Attention Modelling for Videos
- Semantic Compositional Networks for Visual Captioning
- Generating Descriptions with Grounded and Co-Referenced People
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Doubly Attentive Transformer Machine Translation
- An Attempt towards Interpretable Audio-Visual Video Captioning
- Bidirectional Long-Short Term Memory for Video Description
- From Deterministic to Generative: Multi-Modal Stochastic RNNs for Video Captioning
- Video Captioning with Transferred Semantic Attributes
- Video Fill in the Blank with Merging LSTMs
- Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning
- Fully Convolutional Recurrent Network for Handwritten Chinese Text Recognition
- Let's Dance: Learning From Online Dance Videos
- CSVideoNet: A Real-time End-to-end Learning Framework for High-frame-rate Video Compressive Sensing
- Textual Explanations for Self-Driving Vehicles
- MAT: A Multimodal Attentive Translator for Image Captioning
- Road Segmentation Using CNN with GRU
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- Integrating both Visual and Audio Cues for Enhanced Video Caption
- Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
- Learning Visual Storylines with Skipping Recurrent Neural Networks
- Context, Attention and Audio Feature Explorations for Audio Visual Scene-Aware Dialog
- Natural Language Generation with Neural Variational Models
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention
- Title Generation for User Generated Videos
- FASTER Recurrent Networks for Efficient Video Classification
- Multimodal Memory Modelling for Video Captioning
- A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
- Move Forward and Tell: A Progressive Generator of Video Descriptions
- Deep Quantization: Encoding Convolutional Activations with Deep Generative Model
- Video Captioning with Guidance of Multimodal Latent Topics
- Large-scale Video Classification guided by Batch Normalized LSTM Translator
- LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts
- Top-down Visual Saliency Guided by Captions
- Learning long-term dependencies for action recognition with a biologically-inspired deep network
- MAF: Multimodal Alignment Framework for Weakly-Supervised Phrase Grounding
- SparCE: Sparsity aware General Purpose Core Extensions to Accelerate Deep Neural Networks
- DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization
- Action Classification and Highlighting in Videos
- DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Temporal Tessellation: A Unified Approach for Video Analysis
- MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description
- Action Tubelet Detector for Spatio-Temporal Action Localization
- Spatial Memory for Context Reasoning in Object Detection
- Improving Sequential Determinantal Point Processes for Supervised Video Summarization
- Multilevel Language and Vision Integration for Text-to-Clip Retrieval
- Deep Sequence Learning with Auxiliary Information for Traffic Prediction
- Multi-Reference Training with Pseudo-References for Neural Translation and Text Generation
- Learning Spatial-Semantic Context with Fully Convolutional Recurrent Network for Online Handwritten Chinese Text Recognition
- Spatial-Temporal Multi-Cue Network for Continuous Sign Language Recognition
- Y^2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences
- Few-Shot Object Recognition from Machine-Labeled Web Images
- simNet: Stepwise Image-Topic Merging Network for Generating Detailed and Comprehensive Image Captions
- Mining for meaning: from vision to language through multiple networks consensus
- BENCHIP: Benchmarking Intelligence Processors
- Anticipating Daily Intention using On-Wrist Motion Triggered Sensing
- Sequential Person Recognition in Photo Albums with a Recurrent Network
- Semantic Sentence Embeddings for Paraphrasing and Text Summarization
- Multi-modal Dense Video Captioning
- DCA: Diversified Co-Attention towards Informative Live Video Commenting
- Towards Diverse Paragraph Captioning for Untrimmed Videos
- Middle-Out Decoding
- Game-Based Video-Context Dialogue
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Live Video Comment Generation Based on Surrounding Frames and Live Comments
- The Long-Short Story of Movie Description
- mAnI: Movie Amalgamation using Neural Imitation
- Generating Video Descriptions with Topic Guidance
- Weakly Supervised Learning of Heterogeneous Concepts in Videos
- VideoMCC: a New Benchmark for Video Comprehension
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos
- Learning Articulated Motion Models from Visual and Lingual Signals