On the Binding Problem in Artificial Neural Networks
arXiv:2012.05208
Abstract
Contemporary neural networks still fall short of human-level generalization, which extends far beyond our direct experiences. In this paper, we argue that the underlying cause for this shortcoming is their inability to dynamically and flexibly bind information that is distributed throughout the network. This binding problem affects their capacity to acquire a compositional understanding of the world in terms of symbol-like entities (like objects), which is crucial for generalizing in predictable and systematic ways. To address this issue, we propose a unifying framework that revolves around forming meaningful entities from unstructured sensory inputs (segregation), maintaining this separation of information at a representational level (representation), and using these entities to construct new inferences, predictions, and behaviors (composition). Our analysis draws inspiration from a wealth of research in neuroscience and cognitive psychology, and surveys relevant mechanisms from the machine learning literature, to help identify a combination of inductive biases that allow symbolic information processing to emerge naturally in neural networks. We believe that a compositional approach to AI, in terms of grounded symbol-like representations, is of fundamental importance for realizing human-level generalization, and we hope that this paper may contribute towards that goal as a reference and inspiration.
References in corpus (24)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- How transferable are features in deep neural networks?
- Deep Convolutional Networks on Graph-Structured Data
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Recurrent Models of Visual Attention
- Image Segmentation in Video Sequences: A Probabilistic Approach
- Multiple Object Recognition with Visual Attention
- Interaction Networks for Learning about Objects, Relations and Physics
- MONet: Unsupervised Scene Decomposition and Representation
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets
- A Compositional Object-Based Approach to Learning Physical Dynamics
- Using Fast Weights to Attend to the Recent Past
- Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNet
- The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
- A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
- Learning to generalize to new compositions in image understanding
- Independently Controllable Factors
- Object Discovery with a Copy-Pasting GAN
- Stochastic Prediction of Multi-Agent Interactions from Partial Observations
- SCALOR: Generative World Models with Scalable Object Representations
- Towards causal generative scene models via competition of experts
- Continuous Graph Flow