Optimal control as a graphical model inference problem
arXiv:0901.0633 · doi:10.1007/s10994-012-5278-7
Abstract
We reformulate a class of non-linear stochastic optimal control problems introduced by Todorov (2007) as a Kullback-Leibler (KL) minimization problem. As a result, the optimal control computation reduces to an inference computation and approximate inference methods can be applied to efficiently compute approximate optimal controls. We show how this KL control theory contains the path integral control method as a special case. We provide an example of a block stacking task and a multi-agent cooperative game where we demonstrate how approximate inference can be successfully applied to instances that are too complex for exact computation. We discuss the relation of the KL control approach to other inference approaches to control.
26 pages, 12 Figures; Machine Learning Journal (2012)
References in corpus (7)
- Loopy Belief Propagation for Approximate Inference: An Empirical Study
- Expectation Propagation for approximate Bayesian inference
- Optimal control as a graphical model inference problem
- Approximate Inference and Constrained Optimization
- A Method for Using Belief Networks as Influence Diagrams
- Stochastic Optimal Control in Continuous Space-Time Multi-Agent Systems
- KL-learning: Online solution of Kullback-Leibler control problems
Cited by in corpus (92)
- Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review
- Optimal control as a graphical model inference problem
- Thermodynamics as a theory of decision-making with information processing costs
- A free energy principle for a particular physics
- Distral: Robust Multitask Reinforcement Learning
- The free energy principle made simpler but not too simple
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Adaptive importance sampling for control and inference
- Reward Augmented Maximum Likelihood for Neural Structured Prediction
- Efficient Exploration via State Marginal Matching
- Taming the Noise in Reinforcement Learning via Soft Updates
- Reinforcement Learning through Active Inference
- Variational Inverse Control with Events: A General Framework for Data-Driven Reward Definition
- A reinforcement learning approach to rare trajectory sampling
- Data Assimilation: The Schrödinger Perspective
- Designing Ecosystems of Intelligence from First Principles
- The Two Kinds of Free Energy and the Bayesian Revolution
- Bounded rational decision-making from elementary computations that reduce uncertainty
- Neural dynamics under active inference: plausibility and efficiency of information processing
- Reinforcement learning of rare diffusive dynamics
- Generalizing Skills with Semi-Supervised Reinforcement Learning
- Explicit solution of relative entropy weighted control
- Rewriting History with Inverse RL: Hindsight Inference for Policy Improvement
- Information-Theoretic Bounded Rationality
- Equality Constrained Linear Optimal Control With Factor Graphs
- Particle Smoothing for Hidden Diffusion Processes: Adaptive Path Integral Smoother
- Offline Reinforcement Learning with Fisher Divergence Critic Regularization
- A Divergence Minimization Perspective on Imitation Learning Methods
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- Action and Perception as Divergence Minimization
- Intrinsic Motivation and Mental Replay enable Efficient Online Adaptation in Stochastic Recurrent Networks
- Variational approach to rare event simulation using least-squares regression
- Latent Kullback Leibler Control for Continuous-State Systems using Probabilistic Graphical Models
- Approximate Inference and Stochastic Optimal Control
- Stochastic Optimal Control as Non-equilibrium Statistical Mechanics: Calculus of Variations over Density and Current
- Stein Variational Model Predictive Control
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- Black-Box Policy Search with Probabilistic Programs
- On a probabilistic approach to synthesize control policies from example datasets
- An Information-Theoretic Optimality Principle for Deep Reinforcement Learning
- Real-Time Stochastic Optimal Control for Multi-agent Quadrotor Systems
- Probabilistic design of optimal sequential decision-making algorithms in learning and control
- Behavior Priors for Efficient Reinforcement Learning
- A deep active inference model of the rubber-hand illusion
- Making Sense of Reinforcement Learning and Probabilistic Inference
- Real-Time Model Calibration with Deep Reinforcement Learning
- Entropy Regularized Reinforcement Learning Using Large Deviation Theory
- Application of the Free Energy Principle to Estimation and Control
- Geometric Methods for Sampling, Optimisation, Inference and Adaptive Agents
- Free Energy and the Generalized Optimality Equations for Sequential Decision Making
- SOAC: The Soft Option Actor-Critic Architecture
- Optimal Control of Complex Systems through Variational Inference with a Discrete Event Decision Process
- Planning as Inference in Epidemiological Models
- DualSMC: Tunneling Differentiable Filtering and Planning under Continuous POMDPs
- Sparse Randomized Shortest Paths Routing with Tsallis Divergence Regularization
- On Kalman-Bucy filters, linear quadratic control and active inference
- On Reward Function for Survival
- A Constrained Randomized Shortest-Paths Framework for Optimal Exploration
- Hamilton-Jacobi-Bellman Equations for Maximum Entropy Optimal Control
- Active Inference Tree Search in Large POMDPs
- A Multilevel Approach for Stochastic Nonlinear Optimal Control
- Bounded Planning in Passive POMDPs
- Approximately Optimal Continuous-Time Motion Planning and Control via Probabilistic Inference
- Consequential Ranking Algorithms and Long-term Welfare
- Hierarchical Linearly-Solvable Markov Decision Problems
- Forward and Backward Bellman equations improve the efficiency of EM algorithm for DEC-POMDP
- An Information-Theoretic Perspective on Credit Assignment in Reinforcement Learning
- Reward Propagation Using Graph Convolutional Networks
- Globally Optimal Hierarchical Reinforcement Learning for Linearly-Solvable Markov Decision Processes
- Kernel-based diffusion approximated Markov decision processes for autonomous navigation and control on unstructured terrains
- Outcome-Driven Reinforcement Learning via Variational Inference
- A Minimum Relative Entropy Principle for Learning and Acting
- Harmonic Path Integral Diffusion
- Action selection in growing state spaces: Control of Network Structure Growth
- Adaptive Smoothing Path Integral Control
- A Nonparametric Conjugate Prior Distribution for the Maximizing Argument of a Noisy Function
- A conversion between utility and information
- An axiomatic formalization of bounded rationality based on a utility-information equivalence
- Deterministic particle flows for constraining SDEs
- Canonical Cortical Circuits and the Duality of Bayesian Inference and Optimal Control
- Fast Monte-Carlo
- Integration of Imitation Learning using GAIL and Reinforcement Learning using Task-achievement Rewards via Probabilistic Graphical Model
- Joint Perception and Control as Inference with an Object-based Implementation
- Hierarchical model-based policy optimization: from actions to action sequences and back
- Optimal Selective Attention in Reactive Agents
- Deterministic particle flows for constraining stochastic nonlinear systems
- Discrete fully probabilistic design: towards a control pipeline for the synthesis of policies from examples
- Information-Theoretic Stochastic Optimal Control via Incremental Sampling-based Algorithms
- Online Markov decision processes with Kullback-Leibler control cost
- A Minimum Relative Entropy Controller for Undiscounted Markov Decision Processes
- Convergence of Bayesian Control Rule
- Model-Free Risk-Sensitive Reinforcement Learning