The Consciousness Prior
arXiv:1709.08568
Abstract
A new prior is proposed for learning representations of high-level concepts of the kind we manipulate with language. This prior can be combined with other priors in order to help disentangling abstract factors from each other. It is inspired by cognitive neuroscience theories of consciousness, seen as a bottleneck through which just a few elements, after having been selected by attention from a broader pool, are then broadcast and condition further processing, both in perception and decision-making. The set of recently selected elements one becomes aware of is seen as forming a low-dimensional conscious state. This conscious state is combining the few concepts constituting a conscious thought, i.e., what one is immediately conscious of at a particular moment. We claim that this architectural and information-processing constraint corresponds to assumptions about the joint distribution between high-level concepts. To the extent that these assumptions are generally true (and the form of natural language seems consistent with them), they can form a useful prior for representation learning. A low-dimensional thought or conscious state is analogous to a sentence: it involves only a few variables and yet can make a statement with very high probability of being true. This is consistent with a joint distribution (over high-level concepts) which has the form of a sparse factor graph, i.e., where the dependencies captured by each factor of the factor graph involve only very few variables while creating a strong dip in the overall energy function. The consciousness prior also makes it natural to map conscious states to natural language utterances or to express classical AI knowledge in a form similar to facts and rules, albeit capturing uncertainty as well as efficient search mechanisms implemented by attention mechanisms.
References in corpus (2)
Cited by in corpus (57)
- An Introduction to Deep Reinforcement Learning
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
- Weakly-Supervised Disentanglement Without Compromises
- Recurrent Independent Mechanisms
- TDAM: a Topic-Dependent Attention Model for Sentiment Analysis
- Is Attention Better Than Matrix Decomposition?
- On the Binding Problem in Artificial Neural Networks
- Improving Generalization for Abstract Reasoning Tasks Using Disentangled Feature Representations
- Neural Enhanced Belief Propagation on Factor Graphs
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- From Machine Learning to Robotics: Challenges and Opportunities for Embodied Intelligence
- Neural Networks with Recurrent Generative Feedback
- Operationally meaningful representations of physical systems in neural networks
- Nonlinear Invariant Risk Minimization: A Causal Approach
- A visual introduction to Gaussian Belief Propagation
- A Survey on Graph Neural Networks for Knowledge Graph Completion
- Deriving Differential Target Propagation from Iterating Approximate Inverses
- The General Theory of General Intelligence: A Pragmatic Patternist Perspective
- Learning by Abstraction: The Neural State Machine
- An Explicit Local and Global Representation Disentanglement Framework with Applications in Deep Clustering and Unsupervised Object Detection
- Dynamically Pruned Message Passing Networks for Large-Scale Knowledge Graph Reasoning
- A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning
- Learning Awareness Models
- Disentangling Controllable and Uncontrollable Factors of Variation by Interacting with the World
- Latent Causal Invariant Model
- LARNN: Linear Attention Recurrent Neural Network
- Disentanglement via Mechanism Sparsity Regularization: A New Principle for Nonlinear ICA
- ToyArchitecture: Unsupervised Learning of Interpretable Models of the World
- Resolving Spurious Correlations in Causal Models of Environments via Interventions
- An Artificial Consciousness Model and its relations with Philosophy of Mind
- The Tensor Brain: Semantic Decoding for Perception and Memory
- Action-Sufficient State Representation Learning for Control with Structural Constraints
- AliCG: Fine-grained and Evolvable Conceptual Graph Construction for Semantic Search at Alibaba
- Abductive Knowledge Induction From Raw Data
- Deep Learning and the Global Workspace Theory
- State-Denoised Recurrent Neural Networks
- Attention Based Natural Language Grounding by Navigating Virtual Environment
- Deep Anomaly Detection by Residual Adaptation
- Visual Concept Reasoning Networks
- Language (Re)modelling: Towards Embodied Language Understanding
- The neural and cognitive architecture for learning from a small sample
- Compositional Attention: Disentangling Search and Retrieval
- Morphological Computation and Learning to Learn In Natural Intelligent Systems And AI
- Amanuensis: The Programmer's Apprentice
- Neural Consciousness Flow
- Active Observer Visual Problem-Solving Methods are Dynamically Hypothesized, Deployed and Tested
- Theory of Machine Networks: A Case Study
- Image-to-image Mapping with Many Domains by Sparse Attribute Transfer
- Biological Blueprints for Next Generation AI Systems
- State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations
- Hybrid Active Inference
- From internal models toward metacognitive AI
- Using Meta-Knowledge Mined from Identifiers to Improve Intent Recognition in Neuro-Symbolic Algorithms
- Knowledge as Invariance -- History and Perspectives of Knowledge-augmented Machine Learning
- Concepts, Properties and an Approach for Compositional Generalization