Probabilistic design of optimal sequential decision-making algorithms in learning and control
arXiv:2201.05212 · doi:10.1016/j.arcontrol.2022.09.003
Abstract
This survey is focused on certain sequential decision-making problems that involve optimizing over probability functions. We discuss the relevance of these problems for learning and control. The survey is organized around a framework that combines a problem formulation and a set of resolution methods. The formulation consists of an infinite-dimensional optimization problem. The methods come from approaches to search optimal solutions in the space of probability functions. Through the lenses of this overarching framework we revisit popular learning and control algorithms, showing that these naturally arise from suitable variations on the formulation mixed with different resolution methods. A running example, for which we make the code available, complements the survey. Finally, a number of challenges arising from the survey are also outlined.
This is an authors' version of the work that is published in Annual Reviews in Control, Vol. 54, 2022, Pages 81-102. Changes were made to this version by the publisher prior to publication. The final version of record is available at https://doi.org/10.1016/j.arcontrol.2022.09.003
References in corpus (11)
- Conservative Q-Learning for Offline Reinforcement Learning
- Sharing Knowledge in Multi-Task Deep Reinforcement Learning
- Robust Model Predictive Path Integral Control: Analysis and Performance Guarantees
- Stochastic filtering for multiscale stochastic reaction networks based on hybrid approximations
- A Short Survey On Memory Based Reinforcement Learning
- On the crowdsourcing of behaviors for autonomous agents
- Off-Policy Risk Assessment in Contextual Bandits
- On the design of autonomous agents from multiple data sources
- When Is Partially Observable Reinforcement Learning Not Scary?
- Efficient Stochastic Optimal Control through Approximate Bayesian Input Inference
- Control-Tutored Reinforcement Learning: Towards the Integration of Data-Driven and Model-Based Control