Learning Task Specifications from Demonstrations
arXiv:1710.03875
Abstract
Real world applications often naturally decompose into several sub-tasks. In many settings (e.g., robotics) demonstrations provide a natural way to specify the sub-tasks. However, most methods for learning from demonstrations either do not provide guarantees that the artifacts learned for the sub-tasks can be safely recombined or limit the types of composition available. Motivated by this deficit, we consider the problem of inferring Boolean non-Markovian rewards (also known as logical trace properties or specifications) from demonstrations provided by an agent operating in an uncertain, stochastic environment. Crucially, specifications admit well-defined composition rules that are typically easy to interpret. In this paper, we formulate the specification inference task as a maximum a posteriori (MAP) probability inference problem, apply the principle of maximum entropy to derive an analytic demonstration likelihood model and give an efficient approach to search for the most likely specification in a large candidate pool of specifications. In our experiments, we demonstrate how learning specifications can help avoid common problems that often arise due to ad-hoc reward composition.
NIPS 2018
Cited by in corpus (17)
- Towards Verified Artificial Intelligence
- Automatic Discovery of Interpretable Planning Strategies
- Inverse Rational Control with Partially Observable Continuous Nonlinear Dynamics
- Active Finite Reward Automaton Inference and Reinforcement Learning Using Queries and Counterexamples
- A Model Counter's Guide to Probabilistic Systems
- Reactive motion planning with probabilistic safety guarantees
- Graph Temporal Logic Inference for Classification and Identification
- Interactive Robot Training for Non-Markov Tasks
- Learning Branching Heuristics for Propositional Model Counting
- Counter-example Guided Learning of Bounds on Environment Behavior
- Explaining Multi-stage Tasks by Learning Temporal Logic Formulas from Suboptimal Demonstrations
- Learning Finite Linear Temporal Logic Specifications with a Specialized Neural Operator
- Generalized Inverse Planning: Learning Lifted non-Markovian Utility for Generalizable Task Representation
- Adaptive Teaching of Temporal Logic Formulas to Learners with Preferences
- Discretizing Dynamics for Maximum Likelihood Constraint Inference
- Transfer of Temporal Logic Formulas in Reinforcement Learning
- Making Human-Like Trade-offs in Constrained Environments by Learning from Demonstrations