Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
arXiv:1510.07712
Abstract
We present an approach that exploits hierarchical Recurrent Neural Networks (RNNs) to tackle the video captioning problem, i.e., generating one or multiple sentences to describe a realistic video. Our hierarchical framework contains a sentence generator and a paragraph generator. The sentence generator produces one simple short sentence that describes a specific short video interval. It exploits both temporal- and spatial-attention mechanisms to selectively focus on visual elements during generation. The paragraph generator captures the inter-sentence dependency by taking as input the sentential embedding produced by the sentence generator, combining it with the paragraph history, and outputting the new initial state for the sentence generator. We evaluate our approach on two large-scale benchmark datasets: YouTubeClips and TACoS-MultiLevel. The experiments demonstrate that our approach significantly outperforms the current state-of-the-art methods with BLEU@4 scores 0.499 and 0.305 respectively.
In CVPR2016
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Sequence to Sequence Learning with Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
- Sequence Level Training with Recurrent Neural Networks
- Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
- A Hierarchical Neural Autoencoder for Paragraphs and Documents
- Learning a Recurrent Visual Representation for Image Caption Generation
- Question Answering with Subgraph Embeddings
- Jointly Modeling Embedding and Translation to Bridge Video and Language
Cited by in corpus (25)
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- Weakly-Supervised Action Segmentation with Iterative Soft Boundary Assignment
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Word2VisualVec: Image and Video to Sentence Matching by Visual Feature Prediction
- End-to-End Dense Video Captioning with Masked Transformer
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- Weakly Supervised Dense Video Captioning
- Less Is More: Picking Informative Frames for Video Captioning
- V2CNet: A Deep Learning Framework to Translate Videos to Commands for Robotic Manipulation
- Attention-Based Multimodal Fusion for Video Description
- Video Captioning via Hierarchical Reinforcement Learning
- Semantic Compositional Networks for Visual Captioning
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Generating Descriptions with Grounded and Co-Referenced People
- Representation of linguistic form and function in recurrent neural networks
- Self-Gated Memory Recurrent Network for Efficient Scalable HDR Deghosting
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- End-to-End Audio Visual Scene-Aware Dialog using Multimodal Attention-Based Video Features
- Multimodal Memory Modelling for Video Captioning
- Supervising Neural Attention Models for Video Captioning by Human Gaze Data
- MTLE: A Multitask Learning Encoder of Visual Feature Representations for Video and Movie Description
- Improving Classification by Improving Labelling: Introducing Probabilistic Multi-Label Object Interaction Recognition
- Coupled Recurrent Network (CRN)
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos