Uncovering Temporal Context for Video Question and Answering
arXiv:1511.04670
Abstract
In this work, we introduce Video Question Answering in temporal domain to infer the past, describe the present and predict the future. We present an encoder-decoder approach using Recurrent Neural Networks to learn temporal structures of videos and introduce a dual-channel ranking loss to answer multiple-choice questions. We explore approaches for finer understanding of video content using question form of "fill-in-the-blank", and managed to collect 109,895 video clips with duration over 1,000 hours from TACoS, MPII-MD, MEDTest 14 datasets, while the corresponding 390,744 questions are generated from annotations. Extensive experiments demonstrate that our approach significantly outperforms the compared baselines.
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Recurrent Neural Network Regularization
- Unsupervised Learning of Video Representations using LSTMs
- VQA: Visual Question Answering
- Skip-Thought Vectors
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- Visual Madlibs: Fill in the blank Image Generation and Question Answering
- Phrase-based Image Captioning
Cited by in corpus (14)
- A Comprehensive Study of Deep Video Action Recognition
- Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
- Visual Reference Resolution using Attention Memory for Visual Dialog
- Focal Visual-Text Attention for Visual Question Answering
- Learning Language-Visual Embedding for Movie Understanding with Natural-Language
- FVQA: Fact-based Visual Question Answering
- Video Fill in the Blank with Merging LSTMs
- MarioQA: Answering Questions by Watching Gameplay Videos
- A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
- Encoding Video and Label Priors for Multi-label Video Classification on YouTube-8M dataset
- How to Make a BLT Sandwich? Learning to Reason towards Understanding Web Instructional Videos
- VideoMCC: a New Benchmark for Video Comprehension
- The Forgettable-Watcher Model for Video Question Answering
- Bidirectional Multirate Reconstruction for Temporal Modeling in Videos