VGNMN: Video-grounded Neural Module Network to Video-Grounded Language Tasks
arXiv:2104.07921
Abstract
Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue tasks. These tasks extend the complexity of traditional visual tasks with the additional visual temporal variance and language cross-turn dependencies. Motivated by recent NMN approaches on image-grounded tasks, we introduce Video-grounded Neural Module Network (VGNMN) to model the information retrieval process in video-grounded language tasks as a pipeline of neural modules. VGNMN first decomposes all language components in dialogues to explicitly resolve any entity references and detect corresponding action-based inputs from the question. The detected entities and actions are used as parameters to instantiate neural module networks and extract visual cues from the video. Our experiments show that VGNMN can achieve promising performance on a challenging video-grounded dialogue benchmark as well as a video QA benchmark.
Accepted at NAACL 2022 (Oral)
References in corpus (11)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- VQA: Visual Question Answering
- Exploring Models and Data for Image Question Answering
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
- Pathologies of Neural Models Make Interpretations Difficult
- Neural Module Networks for Reasoning over Text
- Bridging Text and Video: A Universal Multimodal Transformer for Video-Audio Scene-Aware Dialog
- Visual Concept-Metaconcept Learning
- Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
- DSTC8-AVSD: Multimodal Semantic Transformer Network with Retrieval Style Word Generator
- Leveraging Topics and Audio Features with Multimodal Attention for Audio Visual Scene-Aware Dialog