Finding Approximate POMDP solutions Through Belief Compression
arXiv:1107.0053 · doi:10.1613/jair.1496
Abstract
Standard value function approaches to finding policies for Partially Observable Markov Decision Processes (POMDPs) are generally considered to be intractable for large models. The intractability of these algorithms is to a large extent a consequence of computing an exact, optimal policy over the entire belief space. However, in real-world POMDP problems, computing the optimal policy for the full belief space is often unnecessary for good control even for problems with complicated policy classes. The beliefs experienced by the controller often lie near a structured, low-dimensional subspace embedded in the high-dimensional belief space. Finding a good approximation to the optimal value function for only this subspace can be much easier than computing the full value function. We introduce a new method for solving large-scale POMDPs by reducing the dimensionality of the belief space. We use Exponential family Principal Components Analysis (Collins, Dasgupta and Schapire, 2002) to represent sparse, high-dimensional belief spaces using small sets of learned features of the belief state. We then plan only in terms of the low-dimensional belief features. By planning in this low-dimensional space, we can find policies for POMDP models that are orders of magnitude larger than models that can be handled by conventional techniques. We demonstrate the use of this algorithm on a synthetic problem and on mobile robot navigation tasks.
References in corpus (13)
- Markov Localization for Mobile Robots in Dynamic Environments
- Value-Function Approximations for Partially Observable Markov Decision Processes
- Tractable Inference for Complex Stochastic Processes
- PEGASUS: A Policy Search Method for Large MDPs and POMDPs
- Finding Approximate POMDP solutions Through Belief Compression
- Stochastic Simulation Algorithms for Dynamic Probabilistic Networks
- Solving POMDPs by Searching in Policy Space
- Learning Finite-State Controllers for Partially Observable Environments
- Solving POMDPs by Searching the Space of Finite Policies
- Speeding Up the Convergence of Value Iteration in Partially Observable Markov Decision Processes
- Reinforcement Learning in POMDP's via Direct Gradient Ascent
- Policy-contingent abstraction for robust robot control
- Learning Geometrically-Constrained Hidden Markov Models for Robot Navigation: Bridging the Topological-Geometrical Gap
Cited by in corpus (17)
- Perseus: Randomized Point-based Value Iteration for POMDPs
- Finding Approximate POMDP solutions Through Belief Compression
- Monte Carlo Sampling Methods for Approximating Interactive POMDPs
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Deep Reinforcement Learning amidst Lifelong Non-Stationarity
- Simplified decision making in the belief space using belief sparsification
- Probabilistic digital twins for geotechnical design and construction
- Poisson noise reduction with non-local PCA
- Exploiting Causality for Selective Belief Filtering in Dynamic Bayesian Networks
- Adaptive Belief Discretization for POMDP Planning
- An Overview of Natural Language State Representation for Reinforcement Learning
- Cooperative multi-agent reinforcement learning for high-dimensional nonequilibrium control
- Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
- How memory architecture affects learning in a simple POMDP: the two-hypothesis testing problem
- Cross-layer estimation and control for Cognitive Radio: Exploiting Sparse Network Dynamics
- On Avoidance Learning with Partial Observability
- Optimal Selective Attention in Reactive Agents