Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
arXiv:1805.00909
Abstract
The framework of reinforcement learning or optimal control provides a mathematical formalization of intelligent decision making that is powerful and broadly applicable. While the general form of the reinforcement learning problem enables effective reasoning about uncertainty, the connection between reinforcement learning and inference in probabilistic models is not immediately obvious. However, such a connection has considerable value when it comes to algorithm design: formalizing a problem as probabilistic inference in principle allows us to bring to bear a wide array of approximate inference tools, extend the model in flexible and powerful ways, and reason about compositionality and partial observability. In this article, we will discuss how a generalization of the reinforcement learning or optimal control problem, which is sometimes termed maximum entropy reinforcement learning, is equivalent to exact probabilistic inference in the case of deterministic dynamics, and variational inference in the case of stochastic dynamics. We will present a detailed derivation of this framework, overview prior work that has drawn on this and related ideas to propose new reinforcement learning and control algorithms, and describe perspectives on future research.
References in corpus (4)
Cited by in corpus (148)
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Exploration in Deep Reinforcement Learning: A Survey
- Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction
- Causal Confusion in Imitation Learning
- The free energy principle made simpler but not too simple
- Game-Theoretic Multiagent Reinforcement Learning
- Trajectory Forecasts in Unknown Environments Conditioned on Grid-Based Plans
- Efficient Exploration via State Marginal Matching
- Dynamics-Aware Unsupervised Discovery of Skills
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
- Reinforcement Learning through Active Inference
- Deep Imitative Models for Flexible Inference, Planning, and Control
- SQIL: Imitation Learning via Reinforcement Learning with Sparse Rewards
- A reinforcement learning approach to rare trajectory sampling
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- The Two Kinds of Free Energy and the Bayesian Revolution
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
- Variational Inference MPC for Bayesian Model-based Reinforcement Learning
- Reinforcement learning of rare diffusive dynamics
- Deconfounding Reinforcement Learning in Observational Settings
- Variational Bayes survival analysis for unemployment modelling
- If MaxEnt RL is the Answer, What is the Question?
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
- VIREL: A Variational Inference Framework for Reinforcement Learning
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- Equality Constrained Linear Optimal Control With Factor Graphs
- A probabilistic generative model for semi-supervised training of coarse-grained surrogates and enforcing physical constraints through virtual observables
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Deep active inference agents using Monte-Carlo methods
- Advances in Variational Inference
- A Divergence Minimization Perspective on Imitation Learning Methods
- Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning
- Active Predictive Coding: A Unified Neural Framework for Learning Hierarchical World Models for Perception and Planning
- Learning Implicit Priors for Motion Optimization
- Action and Perception as Divergence Minimization
- Optimistic Reinforcement Learning by Forward Kullback-Leibler Divergence Optimization
- Boosting Trust Region Policy Optimization by Normalizing Flows Policy
- Unnatural Language Processing: Bridging the Gap Between Synthetic and Natural Language Data
- Stein Variational Model Predictive Control
- Training Agents using Upside-Down Reinforcement Learning
- Evaluating uncertainties in electrochemical impedance spectra of solid oxide fuel cells
- Meta-SAC: Auto-tune the Entropy Temperature of Soft Actor-Critic via Metagradient
- Offline Reinforcement Learning from Images with Latent Space Models
- A Unified Bellman Optimality Principle Combining Reward Maximization and Empowerment
- Probabilistic design of optimal sequential decision-making algorithms in learning and control
- Variational Bayesian Reinforcement Learning with Regret Bounds
- Real-Time Model Calibration with Deep Reinforcement Learning
- Connecting the Dots Between MLE and RL for Sequence Prediction
- Making Sense of Reinforcement Learning and Probabilistic Inference
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
- Scalable Bayesian Inverse Reinforcement Learning
- Powering Hidden Markov Model by Neural Network based Generative Models
- Diverse Auto-Curriculum is Critical for Successful Real-World Multiagent Learning Systems
- Application of the Free Energy Principle to Estimation and Control
- Improving GAN Training with Probability Ratio Clipping and Sample Reweighting
- Towards Stochastic Fault-tolerant Control using Precision Learning and Active Inference
- Disentangled Skill Embeddings for Reinforcement Learning
- PlaNet of the Bayesians: Reconsidering and Improving Deep Planning Network by Incorporating Bayesian Inference
- C-Learning: Learning to Achieve Goals via Recursive Classification
- Greedification Operators for Policy Optimization: Investigating Forward and Reverse KL Divergences
- Grounded Relational Inference: Domain Knowledge Driven Explainable Autonomous Driving
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- Geometric Methods for Sampling, Optimisation, Inference and Adaptive Agents
- Nonlinear Data-Enabled Prediction and Control
- Optimal Control of Complex Systems through Variational Inference with a Discrete Event Decision Process
- SOAC: The Soft Option Actor-Critic Architecture
- Scalable Online Planning via Reinforcement Learning Fine-Tuning
- Differentiable Trust Region Layers for Deep Reinforcement Learning
- Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning
- A Regularized Opponent Model with Maximum Entropy Objective
- Implicit Policy for Reinforcement Learning
- Hierarchical Path-planning from Speech Instructions with Spatial Concept-based Topometric Semantic Mapping
- DualSMC: Tunneling Differentiable Filtering and Planning under Continuous POMDPs
- RLSS: A Deep Reinforcement Learning Algorithm for Sequential Scene Generation
- Low Emission Building Control with Zero-Shot Reinforcement Learning
- Planning as Inference in Epidemiological Models
- Analyzing the Variance of Policy Gradient Estimators for the Linear-Quadratic Regulator
- Exploring the hierarchical structure of human plans via program generation
- Modelling Bounded Rationality in Multi-Agent Interactions by Generalized Recursive Reasoning
- Identifiability in inverse reinforcement learning
- Stochastic Optimal Control as Approximate Input Inference
- An operator view of policy gradient methods
- Optimization Algorithm for Feedback and Feedforward Policies towards Robot Control Robust to Sensing Failures
- Imitation learning for variable speed motion generation over multiple actions
- Model Based Planning with Energy Based Models
- Active Inference or Control as Inference? A Unifying View
- Consequential Ranking Algorithms and Long-term Welfare
- VILD: Variational Imitation Learning with Diverse-quality Demonstrations
- Entropy Regularized Motion Planning via Stein Variational Inference
- Iterative Amortized Policy Optimization
- C-Planning: An Automatic Curriculum for Learning Goal-Reaching Tasks
- f-IRL: Inverse Reinforcement Learning via State Marginal Matching
- Preventing Imitation Learning with Adversarial Policy Ensembles
- Tutorial and Survey on Probabilistic Graphical Model and Variational Inference in Deep Reinforcement Learning
- Outcome-Driven Reinforcement Learning via Variational Inference
- Inverse reinforcement learning for autonomous navigation via differentiable semantic mapping and planning
- An Information-Theoretic Perspective on Credit Assignment in Reinforcement Learning
- Reward Propagation Using Graph Convolutional Networks
- Skill Discovery of Coordination in Multi-agent Reinforcement Learning
- Globally Optimal Hierarchical Reinforcement Learning for Linearly-Solvable Markov Decision Processes
- A Theoretical Connection Between Statistical Physics and Reinforcement Learning
- Control as Hybrid Inference
- Unsupervised Domain Adaptation with Dynamics-Aware Rewards in Reinforcement Learning
- Mutual-Information Regularization in Markov Decision Processes and Actor-Critic Learning
- X2T: Training an X-to-Text Typing Interface with Online Learning from User Feedback
- Domain-Adversarial and Conditional State Space Model for Imitation Learning
- Reinforced Imitation Learning by Free Energy Principle
- Soft Sensors and Process Control using AI and Dynamic Simulation
- Realising Active Inference in Variational Message Passing: the Outcome-blind Certainty Seeker
- Understanding the Origin of Information-Seeking Exploration in Probabilistic Objectives for Control
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- Assisted Perception: Optimizing Observations to Communicate State
- Maximum Entropy Diverse Exploration: Disentangling Maximum Entropy Reinforcement Learning
- On Pathologies in KL-Regularized Reinforcement Learning from Expert Demonstrations
- Reinforcement Learning as Iterative and Amortised Inference
- Online Variational Filtering and Parameter Learning
- Deterministic particle flows for constraining SDEs
- String Diagrams with Factorized Densities
- Neural-to-Tree Policy Distillation with Policy Improvement Criterion
- Inverse Reinforcement Learning via Matching of Optimality Profiles
- Ground States of Quantum Many Body Lattice Models via Reinforcement Learning
- Lyapunov-Based Reinforcement Learning for Decentralized Multi-Agent Control
- Simultaneous estimation of contact position and tool shape with high-dimensional parameters using force measurements and particle filtering
- Reward-Punishment Reinforcement Learning with Maximum Entropy
- Reinforcement Learning for Vision-based Object Manipulation with Non-parametric Policy and Action Primitives
- Self-Paced Deep Reinforcement Learning
- Hierarchical Variational Imitation Learning of Control Programs
- Informing Autonomous Deception Systems with Cyber Expert Performance Data
- Joint Perception and Control as Inference with an Object-based Implementation
- Towards an Understanding of Default Policies in Multitask Policy Optimization
- Global Optimality and Finite Sample Analysis of Softmax Off-Policy Actor Critic under State Distribution Mismatch
- Distilling a Hierarchical Policy for Planning and Control via Representation and Reinforcement Learning
- Improved Reinforcement Learning Coordinated Control of a Mobile Manipulator using Joint Clamping
- Simulation-Based Inference for Global Health Decisions
- Bayesian Inference for Optimal Transport with Stochastic Cost
- Canonical Cortical Circuits and the Duality of Bayesian Inference and Optimal Control
- A Unified View of Algorithms for Path Planning Using Probabilistic Inference on Factor Graphs
- A Probabilistic Interpretation of Self-Paced Learning with Applications to Reinforcement Learning
- Planning and Learning Using Adaptive Entropy Tree Search
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
- Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model
- Weber-Fechner Law in Temporal Difference learning derived from Control as Inference
- Diversity in Action: General-Sum Multi-Agent Continuous Inverse Optimal Control
- Learning What To Do by Simulating the Past
- Model Embedding Model-Based Reinforcement Learning
- Training an Interactive Helper
- Contrastive Active Inference