The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
arXiv:1904.12584
Abstract
We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; instead, our model learns by simply looking at images and reading paired questions and answers. Our model builds an object-based scene representation and translates sentences into executable, symbolic programs. To bridge the learning of two modules, we use a neuro-symbolic reasoning module that executes these programs on the latent scene representation. Analogical to human concept learning, the perception module learns visual concepts based on the language description of the object being referred to. Meanwhile, the learned visual concepts facilitate learning new words and parsing new sentences. We use curriculum learning to guide the searching over the large compositional space of images and language. Extensive experiments demonstrate the accuracy and efficiency of our model on learning visual concepts, word representations, and semantic parsing of sentences. Further, our method allows easy generalization to new object attributes, compositions, language concepts, scenes and questions, and even new program domains. It also empowers applications including visual question answering and bidirectional image-text retrieval.
ICLR 2019 (Oral). Project page: http://nscl.csail.mit.edu/
Cited by in corpus (27)
- Neural-Symbolic Computing: An Effective Methodology for Principled Integration of Machine Learning and Reasoning
- DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering
- On the Binding Problem in Artificial Neural Networks
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Human-in-the-Loop Deep Reinforcement Learning with Application to Autonomous Driving
- ProTo: Program-Guided Transformer for Program-Guided Tasks
- Relational Reasoning using Prior Knowledge for Visual Captioning
- Explainable Artificial Intelligence (XAI) for 6G: Improving Trust between Human and Machine
- Learning from Lexical Perturbations for Consistent Visual Question Answering
- A Benchmark and Baseline for Language-Driven Image Editing
- Supervising the Transfer of Reasoning Patterns in VQA
- From Shallow to Deep Interactions Between Knowledge Representation, Reasoning and Machine Learning (Kay R. Amel group)
- Visual Superordinate Abstraction for Robust Concept Learning
- A Competence-aware Curriculum for Visual Concepts Learning via Question Answering
- Linguistic Structures as Weak Supervision for Visual Scene Graph Generation
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Multi-Granularity Modularized Network for Abstract Visual Reasoning
- Neural Abstructions: Abstractions that Support Construction for Grounded Language Learning
- Neural Abstract Reasoner
- Neuro-Symbolic VQA: A review from the perspective of AGI desiderata
- Web Question Answering with Neurosymbolic Program Synthesis
- Livewired Neural Networks: Making Neurons That Fire Together Wire Together
- Learning by Planning: Language-Guided Global Image Editing
- Domain-robust VQA with diverse datasets and methods but no target labels
- ACRE: Abstract Causal REasoning Beyond Covariation
- Exploring End-to-End Differentiable Natural Logic Modeling
- Extending Answer Set Programs with Neural Networks