Barrier-Certified Adaptive Reinforcement Learning with Applications to Brushbot Navigation
arXiv:1801.09627 · doi:10.1109/TRO.2019.2920206
Abstract
This paper presents a safe learning framework that employs an adaptive model learning algorithm together with barrier certificates for systems with possibly nonstationary agent dynamics. To extract the dynamic structure of the model, we use a sparse optimization technique. We use the learned model in combination with control barrier certificates which constrain policies (feedback controllers) in order to maintain safety, which refers to avoiding particular undesirable regions of the state space. Under certain conditions, recovery of safety in the sense of Lyapunov stability after violations of safety due to the nonstationarity is guaranteed. In addition, we reformulate an action-value function approximation to make any kernel-based nonlinear function estimation method applicable to our adaptive learning framework. Lastly, solutions to the barrier-certified policy optimization are guaranteed to be globally optimal, ensuring the greedy policy improvement under mild conditions. The resulting framework is validated via simulations of a quadrotor, which has previously been used under stationarity assumptions in the safe learnings literature, and is then tested on a real robot, the brushbot, whose dynamics is unknown, highly complex and nonstationary.
©2019 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
References in corpus (6)
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- Constrained Policy Optimization
- Modelling transition dynamics in MDPs with RKHS embeddings
- Hilbert Space Embeddings of POMDPs
- Safe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret
- Online Constrained Model-based Reinforcement Learning
Cited by in corpus (15)
- Adaptive Safety with Control Barrier Functions
- A Survey on Physics Informed Reinforcement Learning: Review and Open Problems
- A Control Barrier Perspective on Episodic Learning via Projection-to-State Safety
- End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks
- Temporal Logic Guided Safe Reinforcement Learning Using Control Barrier Functions
- Learning Hybrid Control Barrier Functions from Data
- Safe Reinforcement Learning via Probabilistic Shields
- Learning Certified Control using Contraction Metric
- Model-based Reinforcement Learning from Signal Temporal Logic Specifications
- Constraint Learning for Control Tasks with Limited Duration Barrier Functions
- Learning Barrier Certificates: Towards Safe Reinforcement Learning with Zero Training-time Violations
- Continuous-time Value Function Approximation in Reproducing Kernel Hilbert Spaces
- Failing with Grace: Learning Neural Network Controllers that are Boundedly Unsafe
- Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC
- A Safety and Passivity Filter for Robot Teleoperation Systems