Building a Large-scale Multimodal Knowledge Base System for Answering Visual Queries
arXiv:1507.05670
Abstract
The complexity of the visual world creates significant challenges for comprehensive visual understanding. In spite of recent successes in visual recognition, today's vision systems would still struggle to deal with visual queries that require a deeper reasoning. We propose a knowledge base (KB) framework to handle an assortment of visual queries, without the need to train new classifiers for new tasks. Building such a large-scale multimodal KB presents a major challenge of scalability. We cast a large-scale MRF into a KB representation, incorporating visual, textual and structured data, as well as their diverse relations. We introduce a scalable knowledge base construction system that is capable of building a KB with half billion variables and millions of parameters in a few hours. Our system achieves competitive results compared to purpose-built models on standard recognition and retrieval tasks, while exhibiting greater flexibility in answering richer visual queries.
References in corpus (5)
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- VQA: Visual Question Answering
- Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
- DimmWitted: A Study of Main-Memory Statistical Analytics
Cited by in corpus (16)
- Explicit Knowledge-based Reasoning for Visual Question Answering
- Ask Me Anything: Free-form Visual Question Answering Based on Knowledge from External Sources
- Visual Question Answering: A Survey of Methods and Datasets
- Visual Affordance and Function Understanding: A Survey
- FVQA: Fact-based Visual Question Answering
- What value do explicit high level concepts have in vision to language problems?
- Iterative Visual Reasoning Beyond Convolutions
- The More You Know: Using Knowledge Graphs for Image Classification
- Learning Visual Knowledge Memory Networks for Visual Question Answering
- Spatial Memory for Context Reasoning in Object Detection
- SCR-Graph: Spatial-Causal Relationships based Graph Reasoning Network for Human Action Prediction
- The Wisdom of MaSSeS: Majority, Subjectivity, and Semantic Similarity in the Evaluation of VQA
- Zero-Shot Scene Graph Relation Prediction through Commonsense Knowledge Integration
- VisualSem: A High-quality Knowledge Graph for Vision and Language
- Cross-Modal Retrieval Augmentation for Multi-Modal Classification
- On the Flip Side: Identifying Counterexamples in Visual Question Answering