Coherent Multi-Sentence Video Description with Variable Level of Detail
arXiv:1403.6173 · doi:10.1007/978-3-319-11752-2_15
Abstract
Humans can easily describe what they see in a coherent way and at varying level of detail. However, existing approaches for automatic video description are mainly focused on single sentence generation and produce descriptions at a fixed level of detail. In this paper, we address both of these limitations: for a variable level of detail we produce coherent multi-sentence descriptions of complex videos. We follow a two-step approach where we first learn to predict a semantic representation (SR) from video and then generate natural language descriptions from the SR. To produce consistent multi-sentence descriptions, we model across-sentence consistency at the level of the SR by enforcing a consistent topic. We also contribute both to the visual recognition of objects proposing a hand-centric approach as well as to the robust generation of sentences using a word lattice. Human judges rate our multi-sentence descriptions as more readable, correct, and relevant than related work. To understand the difference between more detailed and shorter descriptions, we collect and analyze a video description corpus of three levels of detail.
10 pages
References in corpus (1)
Cited by in corpus (51)
- Coherent Multi-Sentence Video Description with Variable Level of Detail
- Recognizing Fine-Grained and Composite Activities using Hand-Centric Features and Script Data
- Hierarchical Boundary-Aware Neural Encoder for Video Captioning
- Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
- Recurrent Topic-Transition GAN for Visual Paragraph Generation
- SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
- A Dataset for Movie Description
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Video Storytelling: Textual Summaries for Events
- Reconstruction Network for Video Captioning
- Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Towards a Visual Turing Challenge
- DeepStory: Video Story QA by Deep Embedded Memory Networks
- Spatio-Temporal Graph for Video Captioning with Knowledge Distillation
- Weakly Supervised Dense Video Captioning
- Auto-captions on GIF: A Large-scale Video-sentence Dataset for Vision-language Pre-training
- Discriminative Latent Semantic Graph for Video Captioning
- Poet: Product-oriented Video Captioner for E-commerce
- Comprehensive Information Integration Modeling Framework for Video Titling
- Video Captioning via Hierarchical Reinforcement Learning
- Controllable Video Captioning with an Exemplar Sentence
- Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval
- End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering
- TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
- Recent Advances and Trends in Multimodal Deep Learning: A Review
- Generating Descriptions with Grounded and Co-Referenced People
- Human-centric Spatio-Temporal Video Grounding With Visual Transformers
- Spatio-Temporal Dynamics and Semantic Attribute Enriched Visual Encoding for Video Captioning
- A Hierarchical Approach for Generating Descriptive Image Paragraphs
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning
- Multimodal Visual Concept Learning with Weakly Supervised Techniques
- Identifying Visible Actions in Lifestyle Vlogs
- Title Generation for User Generated Videos
- LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts
- Reconstruct and Represent Video Contents for Captioning via Reinforcement Learning
- Learning to Compose Topic-Aware Mixture of Experts for Zero-Shot Video Captioning
- Audio-Language Datasets of Scenes and Events: A Survey
- SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
- A Comprehensive Review on Recent Methods and Challenges of Video Description
- MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish
- Video Caption Dataset for Describing Human Actions in Japanese
- DVCFlow: Modeling Information Flow Towards Human-like Video Captioning
- Empirical Autopsy of Deep Video Captioning Frameworks
- Semantic Sentence Embeddings for Paraphrasing and Text Summarization
- DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
- Stories in the Eye: Contextual Visual Interactions for Efficient Video to Language Translation
- Narration Generation for Cartoon Videos
- Multi-modal Dense Video Captioning
- Syntax Customized Video Captioning by Imitating Exemplar Sentences
- The Long-Short Story of Movie Description