TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
arXiv:1704.04497
Abstract
Vision and language understanding has emerged as a subject undergoing intense study in Artificial Intelligence. Among many tasks in this line of research, visual question answering (VQA) has been one of the most successful ones, where the goal is to learn a model that understands visual content at region-level details and finds their associations with pairs of questions and answers in the natural language form. Despite the rapid progress in the past few years, most existing work in VQA have focused primarily on images. In this paper, we focus on extending VQA to the video domain and contribute to the literature in three important ways. First, we propose three new tasks designed specifically for video VQA, which require spatio-temporal reasoning from videos to answer questions correctly. Next, we introduce a new large-scale dataset for video VQA named TGIF-QA that extends existing VQA work with our new tasks. Finally, we propose a dual-LSTM based approach with both spatial and temporal attention, and show its effectiveness over conventional VQA techniques through empirical evaluations.
Accepted paper at CVPR 2017 (Spotlight)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Layer Normalization
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- Image Captioning with Semantic Attention
- TGIF: A New Dataset and Benchmark on Animated GIF Description
- Video2GIF: Automatic Generation of Animated GIFs from Video
Cited by in corpus (15)
- Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering
- TVQA+: Spatio-Temporal Grounding for Video Question Answering
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Polysemous Visual-Semantic Embedding for Cross-Modal Retrieval
- Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
- A Read-Write Memory Network for Movie Story Understanding
- Character Matters: Video Story Understanding with Character-Aware Relations
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QA
- Cross-Modal Retrieval with Implicit Concept Association
- Constructing Hierarchical Q&A Datasets for Video Story Understanding
- Efficient Video Classification Using Fewer Frames
- Location-aware Graph Convolutional Networks for Video Question Answering
- I Have Seen Enough: A Teacher Student Network for Video Classification Using Fewer Frames
- CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising