A Read-Write Memory Network for Movie Story Understanding
arXiv:1709.09345
Abstract
We propose a novel memory network model named Read-Write Memory Network (RWMN) to perform question and answering tasks for large-scale, multimodal movie story understanding. The key focus of our RWMN model is to design the read network and the write network that consist of multiple convolutional layers, which enable memory read and write operations to have high capacity and flexibility. While existing memory-augmented network models treat each memory slot as an independent block, our use of multi-layered CNNs allows the model to read and write sequential memory cells as chunks, which is more reasonable to represent a sequential story because adjacent memory blocks often have strong correlations. For evaluation, we apply our model to all the six tasks of the MovieQA benchmark, and achieve the best accuracies on several tasks, especially on the visual QA task. Our model shows a potential to better understand not only the content in the story, but also more abstract information, such as relationships between characters and the reasons for their actions.
accepted paper at ICCV 2017
References in corpus (8)
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
- Hierarchical Memory Networks
- TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- DeepStory: Video Story QA by Deep Embedded Memory Networks
- End-to-end Concept Word Detection for Video Captioning, Retrieval, and Question Answering
Cited by in corpus (7)
- Motion-Appearance Co-Memory Networks for Video Question Answering
- Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
- Character Matters: Video Story Understanding with Character-Aware Relations
- ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
- A Memory Network Approach for Story-based Temporal Summarization of 360° Videos
- Speaker Naming in Movies
- Adversarial Multimodal Network for Movie Question Answering