On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models
arXiv:1511.09249
Abstract
This paper addresses the general problem of reinforcement learning (RL) in partially observable environments. In 2013, our large RL recurrent neural networks (RNNs) learned from scratch to drive simulated cars from high-dimensional video input. However, real brains are more powerful in many ways. In particular, they learn a predictive model of their initially unknown environment, and somehow use it for abstract (e.g., hierarchical) planning and reasoning. Guided by algorithmic information theory, we describe RNN-based AIs (RNNAIs) designed to do the same. Such an RNNAI can be trained on never-ending sequences of tasks, some of them provided by the user, others invented by the RNNAI itself in a curious, playful fashion, to improve its RNN-based world model. Unlike our previous model-building RNN-based RL machines dating back to 1990, the RNNAI learns to actively query its model for abstract reasoning and planning and decision making, essentially "learning to think." The basic ideas of this report can be applied to many other cases where one RNN-like system exploits the algorithmic information content of another. They are taken from a grant proposal submitted in Fall 2014, and also explain concepts such as "mirror neurons." Experimental results will be described in separate papers.
36 pages, 1 figure. arXiv admin note: substantial text overlap with arXiv:1404.7828
References in corpus (10)
- Deep Learning in Neural Networks: An Overview
- Sequence to Sequence Learning with Neural Networks
- Playing Atari with Deep Reinforcement Learning
- Going Deeper with Convolutions
- Multi-digit Number Recognition from Street View Imagery using Deep Convolutional Neural Networks
- Visualizing and Understanding Convolutional Networks
- A Clockwork RNN
- Show and Tell: A Neural Image Caption Generator
- Multi-column Deep Neural Networks for Image Classification
- Self-Delimiting Neural Networks
Cited by in corpus (22)
- Unity: A General Platform for Intelligent Agents
- Learning a Driving Simulator
- The Predictron: End-To-End Learning and Planning
- Relational Neural Expectation Maximization: Unsupervised Discovery of Objects and their Interactions
- Artificial Intelligence and its Role in Near Future
- Some Considerations on Learning to Explore via Meta-Reinforcement Learning
- Will we ever have Conscious Machines?
- A Comprehensive Overview and Survey of Recent Advances in Meta-Learning
- A Survey of Deep Learning Techniques for Mobile Robot Applications
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEs
- Predictive Information Accelerates Learning in RL
- On the role of planning in model-based deep reinforcement learning
- Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
- Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
- Jointly-Learned State-Action Embedding for Efficient Reinforcement Learning
- One Big Net For Everything
- Continual Learning Using World Models for Pseudo-Rehearsal
- Derivatives of Turing machines in Linear Logic
- Generative Adversarial Networks are Special Cases of Artificial Curiosity (1990) and also Closely Related to Predictability Minimization (1991)
- Deep Reinforcement Learning From Raw Pixels in Doom
- A Brief Survey of Associations Between Meta-Learning and General AI
- Learning Shared Dynamics with Meta-World Models