The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
arXiv:1904.12584
Abstract
We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; instead, our model learns by simply looking at images and reading paired questions and answers. Our model builds an object-based scene representation and translates sentences into executable, symbolic programs. To bridge the learning of two modules, we use a neuro-symbolic reasoning module that executes these programs on the latent scene representation. Analogical to human concept learning, the perception module learns visual concepts based on the language description of the object being referred to. Meanwhile, the learned visual concepts facilitate learning new words and parsing new sentences. We use curriculum learning to guide the searching over the large compositional space of images and language. Extensive experiments demonstrate the accuracy and efficiency of our model on learning visual concepts, word representations, and semantic parsing of sentences. Further, our method allows easy generalization to new object attributes, compositions, language concepts, scenes and questions, and even new program domains. It also empowers applications including visual question answering and bidirectional image-text retrieval.
ICLR 2019 (Oral). Project page: http://nscl.csail.mit.edu/
Cited by in corpus (26)
- Neural-Symbolic Computing: An Effective Methodology for Principled Integration of Machine Learning and Reasoning
- DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
- On the Binding Problem in Artificial Neural Networks
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Human-in-the-Loop Deep Reinforcement Learning with Application to Autonomous Driving
- ProTo: Program-Guided Transformer for Program-Guided Tasks
- Relational Reasoning using Prior Knowledge for Visual Captioning
- Learning from Lexical Perturbations for Consistent Visual Question Answering
- Explainable Artificial Intelligence (XAI) for 6G: Improving Trust between Human and Machine
- Supervising the Transfer of Reasoning Patterns in VQA
- From Shallow to Deep Interactions Between Knowledge Representation, Reasoning and Machine Learning (Kay R. Amel group)
- A Benchmark and Baseline for Language-Driven Image Editing
- A Competence-aware Curriculum for Visual Concepts Learning via Question Answering
- Linguistic Structures as Weak Supervision for Visual Scene Graph Generation
- Multi-Granularity Modularized Network for Abstract Visual Reasoning
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- Web Question Answering with Neurosymbolic Program Synthesis
- Neuro-Symbolic VQA: A review from the perspective of AGI desiderata
- Neural Abstract Reasoner
- Neural Abstructions: Abstractions that Support Construction for Grounded Language Learning
- Exploring End-to-End Differentiable Natural Logic Modeling
- Learning by Planning: Language-Guided Global Image Editing
- ACRE: Abstract Causal REasoning Beyond Covariation
- Domain-robust VQA with diverse datasets and methods but no target labels
- Extending Answer Set Programs with Neural Networks