Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning
arXiv:1509.08731
Abstract
The mutual information is a core statistical quantity that has applications in all areas of machine learning, whether this is in training of density models over multiple data modalities, in maximising the efficiency of noisy transmission channels, or when learning behaviour policies for exploration by artificial agents. Most learning algorithms that involve optimisation of the mutual information rely on the Blahut-Arimoto algorithm --- an enumerative algorithm with exponential complexity that is not suitable for modern machine learning applications. This paper provides a new approach for scalable optimisation of the mutual information by merging techniques from variational inference and deep learning. We develop our approach by focusing on the problem of intrinsically-motivated learning, where the mutual information forms the definition of a well-known internal drive known as empowerment. Using a variational lower bound on the mutual information, combined with convolutional networks for handling visual input streams, we develop a stochastic optimisation algorithm that allows for scalable information maximisation and empowerment-based reasoning directly from pixels to actions.
Proceedings of the 29th Conference on Neural Information Processing Systems (NIPS 2015)
References in corpus (1)
Cited by in corpus (99)
- A Brief Survey of Deep Reinforcement Learning
- An Introduction to Variational Autoencoders
- On the Opportunities and Risks of Foundation Models
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- Exploration in Deep Reinforcement Learning: A Survey
- VIME: Variational Information Maximizing Exploration
- Emergent Tool Use From Multi-Agent Autocurricula
- Large-Scale Study of Curiosity-Driven Learning
- Deep Variational Information Bottleneck
- Unifying Count-Based Exploration and Intrinsic Motivation
- Variational Intrinsic Control
- Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- A survey on intrinsic motivation in reinforcement learning
- Deep Successor Reinforcement Learning
- Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
- Stein Variational Policy Gradient
- Dynamics-Aware Unsupervised Discovery of Skills
- Intrinsically motivated reinforcement learning for human-robot interaction in the real-world
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- Skew-Fit: State-Covering Self-Supervised Reinforcement Learning
- Reinforcement Learning through Active Inference
- Learning Likelihoods with Conditional Normalizing Flows
- Intrinsically motivated collective motion
- Variational Information Maximization for Feature Selection
- Curiosity-Driven Experience Prioritization via Density Estimation
- Information-Theoretic Bounded Rationality
- Learning to Reach Goals via Iterated Supervised Learning
- A Unified Bellman Equation for Causal Information and Value in Markov Decision Processes
- Deep Reinforcement and InfoMax Learning
- Action and Perception as Divergence Minimization
- Penalizing side effects using stepwise relative reachability
- Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning
- Fast Task Inference with Variational Intrinsic Successor Features
- BeBold: Exploration Beyond the Boundary of Explored Regions
- Structure in Deep Reinforcement Learning: A Survey and Open Problems
- A Unified Bellman Optimality Principle Combining Reward Maximization and Empowerment
- Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Conservative Agency via Attainable Utility Preservation
- AvE: Assistance via Empowerment
- Curiosity-driven reinforcement learning with homeostatic regulation
- Learning Goal Embeddings via Self-Play for Hierarchical Reinforcement Learning
- SMiRL: Surprise Minimizing Reinforcement Learning in Unstable Environments
- Beyond Fine-Tuning: Transferring Behavior in Reinforcement Learning
- Learning Affordance Landscapes for Interaction Exploration in 3D Environments
- Behavior From the Void: Unsupervised Active Pre-Training
- Deep Reinforcement Learning for Clinical Decision Support: A Brief Survey
- Curiosity-Driven Multi-Criteria Hindsight Experience Replay
- MADE: Exploration via Maximizing Deviation from Explored Regions
- Evaluating Agents without Rewards
- Adaptive Reward-Free Exploration
- Ensemble Estimation of Generalized Mutual Information with Applications to Genomics
- Learning Awareness Models
- Causal Influence Detection for Improving Efficiency in Reinforcement Learning
- Information Theoretically Aided Reinforcement Learning for Embodied Agents
- Ready Policy One: World Building Through Active Learning
- The Variational Bandwidth Bottleneck: Stochastic Evaluation on an Information Budget
- Novelty Search in Representational Space for Sample Efficient Exploration
- Is Curiosity All You Need? On the Utility of Emergent Behaviours from Curious Exploration
- Adversarial Imitation via Variational Inverse Reinforcement Learning
- Co-GAIL: Learning Diverse Strategies for Human-Robot Collaboration
- Learning Efficient Representation for Intrinsic Motivation
- Active World Model Learning with Progress Curiosity
- Influence-Based Multi-Agent Exploration
- Task-Agnostic Morphology Evolution
- Implicit Generative Modeling for Efficient Exploration
- Maximizing Information Gain in Partially Observable Environments via Prediction Reward
- Large scale evaluation of importance maps in automatic speech recognition
- Reinforcement Learning-based Visual Navigation with Information-Theoretic Regularization
- Variational Empowerment as Representation Learning for Goal-Based Reinforcement Learning
- Empowerment-driven Exploration using Mutual Information Estimation
- New And Surprising Ways to Be Mean. Adversarial NPCs with Coupled Empowerment Minimisation
- Mutual Information State Intrinsic Control
- Simple Sensor Intentions for Exploration
- Skill Discovery of Coordination in Multi-agent Reinforcement Learning
- The Journey is the Reward: Unsupervised Learning of Influential Trajectories
- The Information Geometry of Unsupervised Reinforcement Learning
- Balancing New Against Old Information: The Role of Surprise in Learning
- Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework
- Maximum Entropy Diverse Exploration: Disentangling Maximum Entropy Reinforcement Learning
- Efficient Empowerment Estimation for Unsupervised Stabilization
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- An Empowerment-based Solution to Robotic Manipulation Tasks with Sparse Rewards
- The Barbados 2018 List of Open Issues in Continual Learning
- Mutual Information-based State-Control for Intrinsically Motivated Reinforcement Learning
- Understanding the Origin of Information-Seeking Exploration in Probabilistic Objectives for Control
- Don't Do What Doesn't Matter: Intrinsic Motivation with Action Usefulness
- Neurons Activation Visualization and Information Theoretic Analysis
- Learning in Sparse Rewards settings through Quality-Diversity algorithms
- Autonomous sPOMDP Environment Modeling With Partial Model Exploitation
- Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching
- Learning Time-Sensitive Strategies in Space Fortress
- Hard Attention Control By Mutual Information Maximization
- Progressive growing of self-organized hierarchical representations for exploration
- Variational Intrinsic Control Revisited
- Option Discovery in the Absence of Rewards with Manifold Analysis
- Experimental Evidence that Empowerment May Drive Exploration in Sparse-Reward Environments
- Unsupervised Skill-Discovery and Skill-Learning in Minecraft