Learning Finite-State Controllers for Partially Observable Environments
arXiv:1301.6721
Abstract
Reactive (memoryless) policies are sufficient in completely observable Markov decision processes (MDPs), but some kind of memory is usually necessary for optimal control of a partially observable MDP. Policies with finite memory can be represented as finite-state automata. In this paper, we extend Baird and Moore's VAPS algorithm to the problem of learning general finite-state automata. Because it performs stochastic gradient descent, this algorithm can be shown to converge to a locally optimal finite-state controller. We provide the details of the algorithm and then consider the question of under what conditions stochastic gradient descent will outperform exact gradient descent. We conclude with empirical results comparing the performance of stochastic and exact gradient descent, and showing the ability of our algorithm to extract the useful information contained in the sequence of past observations to compensate for the lack of observability at each time-step.
Appears in Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence (UAI1999)
References in corpus (3)
Cited by in corpus (27)
- Deep Learning in Neural Networks: An Overview
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- PEGASUS: A Policy Search Method for Large MDPs and POMDPs
- Learning to Cooperate via Policy Search
- Solving POMDPs by Searching the Space of Finite Policies
- Learning Partially Observable Deterministic Action Models
- Nonapproximability Results for Partially Observable Markov Decision Processes
- Hierarchical POMDP Controller Optimization by Likelihood Maximization
- On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models
- Inverse Reinforcement Learning in Swarm Systems
- The Thing That We Tried Didn't Work Very Well : Deictic Representation in Reinforcement Learning
- Induction and Exploitation of Subgoal Automata for Reinforcement Learning
- Approximate information state for approximate planning and reinforcement learning in partially observed systems
- Memory-Limited Partially Observable Stochastic Control and its Mean-Field Control Approach
- A Sufficient Statistic for Influence in Structured Multiagent Environments
- A Bayesian Approach to Policy Recognition and State Representation Learning
- PAC Reinforcement Learning with Rich Observations
- Towards Understanding Cooperative Multi-Agent Q-Learning with Value Factorization
- Learning Deep Neural Network Policies with Continuous Memory States
- Sparse Stochastic Finite-State Controllers for POMDPs
- Learning What to Memorize: Using Intrinsic Motivation to Form Useful Memory in Partially Observable Reinforcement Learning
- Team Behavior in Interactive Dynamic Influence Diagrams with Applications to Ad Hoc Teams
- Reconciling Rewards with Predictive State Representations
- A Function Approximation Approach to Estimation of Policy Gradient for POMDP with Structured Policies
- A Gauss-Newton Method for Markov Decision Processes
- How memory architecture affects learning in a simple POMDP: the two-hypothesis testing problem
- Cross-Entropic Learning of a Machine for the Decision in a Partially Observable Universe