Safe Model-based Reinforcement Learning with Stability Guarantees
arXiv:1705.08551
Abstract
Reinforcement learning is a powerful paradigm for learning optimal policies from experimental data. However, to find optimal policies, most reinforcement learning algorithms explore all possible actions, which may be harmful for real-world systems. As a consequence, learning algorithms are rarely applied on safety-critical systems in the real world. In this paper, we present a learning algorithm that explicitly considers safety, defined in terms of stability guarantees. Specifically, we extend control-theoretic results on Lyapunov stability verification and show how to use statistical models of the dynamics to obtain high-performance control policies with provable stability certificates. Moreover, under additional regularity assumptions in terms of a Gaussian process prior, we prove that one can effectively and safely collect data in order to learn about the dynamics and thus both improve control performance and expand the safe region of the state space. In our experiments, we show how the resulting algorithm can safely optimize a neural network policy on a simulated inverted pendulum, without the pendulum ever falling down.
Proc. of Neural Information Processing Systems (NIPS), 2017
Cited by in corpus (120)
- On the Opportunities and Risks of Foundation Models
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Safe Reinforcement Learning Using Robust MPC
- Neural Lander: Stable Drone Landing Control using Learned Dynamics
- Verifiable Reinforcement Learning via Policy Extraction
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Contraction Theory for Nonlinear Stability Analysis and Learning-based Control: A Tutorial Overview
- Online learning-based Model Predictive Control with Gaussian Process Models and Stability Guarantees
- Explainable Deep One-Class Classification
- Learning for Safety-Critical Control with Control Barrier Functions
- Learning to Walk in the Real World with Minimal Human Effort
- Bayesian Learning-Based Adaptive Control for Safety Critical Systems
- Active Learning in Robotics: A Review of Control Principles
- Learning Control Lyapunov Functions from Counterexamples and Demonstrations
- Data-based stabilization of unknown bilinear systems with guaranteed basin of attraction
- Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning
- ACES -- Automatic Configuration of Energy Harvesting Sensors with Reinforcement Learning
- On the design of terminal ingredients for data-driven MPC
- Efficient and Accurate Estimation of Lipschitz Constants for Deep Neural Networks
- Certified Adversarial Robustness for Deep Reinforcement Learning
- Safe Reinforcement Learning via Curriculum Induction
- Enforcing Policy Feasibility Constraints through Differentiable Projection for Energy Optimization
- Conservative Safety Critics for Exploration
- Robust Control for Dynamical Systems With Non-Gaussian Noise via Formal Abstractions
- Interpretable PID Parameter Tuning for Control Engineering using General Dynamic Neural Networks: An Extensive Comparison
- Exploration-Exploitation in Constrained MDPs
- Model-Based Policy Search Using Monte Carlo Gradient Estimation with Real Systems Application
- Lipschitz constant estimation of Neural Networks via sparse polynomial optimization
- Constrained Upper Confidence Reinforcement Learning
- Combating the Compounding-Error Problem with a Multi-step Model
- An Introduction to Gaussian Process Models
- Sampling-Based Robust Control of Autonomous Systems with Non-Gaussian Noise
- MAMPS: Safe Multi-Agent Reinforcement Learning via Model Predictive Shielding
- Recovery RL: Safe Reinforcement Learning with Learned Recovery Zones
- Safe Multi-Agent Interaction through Robust Control Barrier Functions with Learned Uncertainties
- PAC Confidence Sets for Deep Neural Networks via Calibrated Prediction
- SAMBA: Safe Model-Based & Active Reinforcement Learning
- Safety-Guided Deep Reinforcement Learning via Online Gaussian Process Estimation
- Uniform Error Bounds for Gaussian Process Regression with Application to Safe Control
- Strategy Synthesis for Partially-known Switched Stochastic Systems
- Conservative Agency via Attainable Utility Preservation
- Efficient Learning of a Linear Dynamical System with Stability Guarantees
- Constrained Model-based Reinforcement Learning with Robust Cross-Entropy Method
- Learning-based Model Predictive Control for Safe Exploration
- Model-based Reinforcement Learning from Signal Temporal Logic Specifications
- Neurosymbolic Reinforcement Learning with Formally Verified Exploration
- Safe Learning-based Gradient-free Model Predictive Control Based on Cross-entropy Method
- Variable impedance control and learning -- A review
- Enforcing robust control guarantees within neural network policies
- Off-Dynamics Reinforcement Learning: Training for Transfer with Domain Classifiers
- Uniform Error and Posterior Variance Bounds for Gaussian Process Regression with Application to Safe Control
- Safe Interactive Model-Based Learning
- Posterior Variance Analysis of Gaussian Processes with Application to Average Learning Curves
- Safe Control Algorithms Using Energy Functions: A Unified Framework, Benchmark, and New Directions
- Lyapunov Barrier Policy Optimization
- Almost Surely Stable Deep Dynamics
- Safety Augmented Value Estimation from Demonstrations (SAVED): Safe Deep Model-Based RL for Sparse Cost Robotic Tasks
- Imitation-Projected Programmatic Reinforcement Learning
- Value Iteration in Continuous Actions, States and Time
- Conservative Exploration in Reinforcement Learning
- Robust Model Predictive Shielding for Safe Reinforcement Learning with Stochastic Dynamics
- Pseudo-Convolutional Policy Gradient for Sequence-to-Sequence Lip-Reading
- Robust Regression for Safe Exploration in Control
- Safe Reinforcement Learning with Natural Language Constraints
- Practical Reinforcement Learning For MPC: Learning from sparse objectives in under an hour on a real robot
- Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations
- MRAC-RL: A Framework for On-Line Policy Adaptation Under Parametric Model Uncertainty
- Safe Exploration in Markov Decision Processes with Time-Variant Safety using Spatio-Temporal Gaussian Process
- Provably Correct Training of Neural Network Controllers Using Reachability Analysis
- Online Learning in Kernelized Markov Decision Processes
- Robust exploration in linear quadratic reinforcement learning
- Provably Safe PAC-MDP Exploration Using Analogies
- Evaluating the progress of Deep Reinforcement Learning in the real world: aligning domain-agnostic and domain-specific research
- Learning Safe Neural Network Controllers with Barrier Certificates
- Certainty Equivalent Perception-Based Control
- Tutoring Reinforcement Learning via Feedback Control
- Avoiding Side Effects in Complex Environments
- Explicit Explore, Exploit, or Escape (): near-optimal safety-constrained reinforcement learning in polynomial time
- Control-Tutored Reinforcement Learning
- Learning Nonlinear State Space Models with Hamiltonian Sequential Monte Carlo Sampler
- Stable Reinforcement Learning with Unbounded State Space
- Reinforcement Learning Control of Constrained Dynamic Systems with Uniformly Ultimate Boundedness Stability Guarantee
- Neural Lyapunov Redesign
- Gaussian Process Uniform Error Bounds with Unknown Hyperparameters for Safety-Critical Applications
- Uncertainty-aware Safe Exploratory Planning using Gaussian Process and Neural Control Contraction Metric
- Safe Reinforcement Learning with Linear Function Approximation
- Safety Verification of Model Based Reinforcement Learning Controllers
- ALPaCA vs. GP-based Prior Learning: A Comparison between two Bayesian Meta-Learning Algorithms
- Learning Dynamics Models with Stable Invariant Sets
- Safely Learning Dynamical Systems from Short Trajectories
- Fully Bayesian Recurrent Neural Networks for Safe Reinforcement Learning
- Auditing Robot Learning for Safety and Compliance during Deployment
- Stabilizing Neural Control Using Self-Learned Almost Lyapunov Critics
- Representation of Reinforcement Learning Policies in Reproducing Kernel Hilbert Spaces
- DESTA: A Framework for Safe Reinforcement Learning with Markov Games of Intervention
- Probabilistic Safety for Bayesian Neural Networks
- Learning to Be Cautious
- Enhancement for Robustness of Koopman Operator-based Data-driven Mobile Robotic Systems
- Exact Asymptotics for Linear Quadratic Adaptive Control
- Formal controller synthesis for hybrid systems using genetic programming
- Safe Distributional Reinforcement Learning
- In Proximity of ReLU DNN, PWA Function, and Explicit MPC
- Sensitivity and safety of fully probabilistic control
- Protective Policy Transfer
- Online Algorithms and Policies Using Adaptive and Machine Learning Approaches
- Constrained Policy Gradient Method for Safe and Fast Reinforcement Learning: a Neural Tangent Kernel Based Approach
- -: Adaptive Control with Bayesian Learning
- Efficient Reinforcement Learning in Resource Allocation Problems Through Permutation Invariant Multi-task Learning
- Automatic Exploration Process Adjustment for Safe Reinforcement Learning with Joint Chance Constraint Satisfaction
- FISAR: Forward Invariant Safe Reinforcement Learning with a Deep Neural Network-Based Optimize
- Variance-Based Risk Estimations in Markov Processes via Transformation with State Lumping
- Contraction -Adaptive Control using Gaussian Processes
- Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems
- Autonomous Reinforcement Learning via Subgoal Curricula
- Youla-REN: Learning Nonlinear Feedback Policies with Robust Stability Guarantees
- Is the Rush to Machine Learning Jeopardizing Safety? Results of a Survey
- Safe Control of Arbitrary Nonlinear Systems using Dynamic Extension
- Learning Deep Energy Shaping Policies for Stability-Guaranteed Manipulation
- Can a Compact Neuronal Circuit Policy be Re-purposed to Learn Simple Robotic Control?