Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning
arXiv:1909.07299 · doi:10.1109/ICRA40945.2020.9196796
Abstract
We present a reinforcement learning (RL) framework to synthesize a control policy from a given linear temporal logic (LTL) specification in an unknown stochastic environment that can be modeled as a Markov Decision Process (MDP). Specifically, we learn a policy that maximizes the probability of satisfying the LTL formula without learning the transition probabilities. We introduce a novel rewarding and path-dependent discounting mechanism based on the LTL formula such that (i) an optimal policy maximizing the total discounted reward effectively maximizes the probabilities of satisfying LTL objectives, and (ii) a model-free RL algorithm using these rewards and discount factors is guaranteed to converge to such policy. Finally, we illustrate the applicability of our RL-based synthesis approach on two motion planning case studies.
Cited by in corpus (6)
- Reward Machines: Exploiting Reward Function Structure in Reinforcement Learning
- Reinforcement Learning of Control Policy for Linear Temporal Logic Specifications Using Limit-Deterministic Generalized Büchi Automata
- Model-Free Reinforcement Learning for Stochastic Games with Linear Temporal Logic Objectives
- Reinforcement Learning Based Temporal Logic Control with Soft Constraints Using Limit-deterministic Generalized Buchi Automata
- Synthesis of Discounted-Reward Optimal Policies for Markov Decision Processes Under Linear Temporal Logic Specifications
- Reward Shaping for Reinforcement Learning with Omega-Regular Objectives