Perseus: Randomized Point-based Value Iteration for POMDPs
arXiv:1109.2145 · doi:10.1613/jair.1659
Abstract
Partially observable Markov decision processes (POMDPs) form an attractive and principled framework for agent planning under uncertainty. Point-based approximate techniques for POMDPs compute a policy based on a finite set of points collected in advance from the agents belief space. We present a randomized point-based value iteration algorithm called Perseus. The algorithm performs approximate value backup stages, ensuring that in each backup stage the value of each point in the belief set is improved; the key observation is that a single backup may improve the value of many belief points. Contrary to other point-based methods, Perseus backs up only a (randomly selected) subset of points in the belief set, sufficient for improving the value of each belief point in the set. We show how the same idea can be extended to dealing with continuous action spaces. Experimental results show the potential of Perseus in large scale POMDP problems.
References in corpus (5)
- Value-Function Approximations for Partially Observable Markov Decision Processes
- Incremental Pruning: A Simple, Fast, Exact Method for Partially Observable Markov Decision Processes
- Finding Approximate POMDP solutions Through Belief Compression
- Solving POMDPs by Searching the Space of Finite Policies
- Speeding Up the Convergence of Value Iteration in Partially Observable Markov Decision Processes
Cited by in corpus (19)
- Partially Observable Markov Decision Processes in Robotics: A Survey
- Partially Observable Markov Decision Processes (POMDPs) and Robotics
- Planning for robotic exploration based on forward simulation
- Optimal Inspection and Maintenance Planning for Deteriorating Structural Components through Dynamic Bayesian Networks and Markov Decision Processes
- Quantum POMDPs
- Inference and dynamic decision-making for deteriorating systems with probabilistic dependencies through Bayesian networks and deep reinforcement learning
- A distributed, plug-n-play algorithm for multi-robot applications with a priori non-computable objective functions
- Optimal active particle navigation meets machine learning
- Perception-aware Autonomous Mast Motion Planning for Planetary Exploration Rovers
- Deep reinforcement learning for the olfactory search POMDP: a quantitative benchmark
- Knowledge Transfer for Cross-Domain Reinforcement Learning: A Systematic Review
- Exploiting Submodular Value Functions For Scaling Up Active Perception
- Optimal policies for Bayesian olfactory search in turbulent flows
- Knowledge-Based Hierarchical POMDPs for Task Planning
- Probabilistic design of optimal sequential decision-making algorithms in learning and control
- HARPS: An Online POMDP Framework for Human-Assisted Robotic Planning and Sensing
- Active Inference Tree Search in Large POMDPs
- Optimal Sensing via Multi-armed Bandit Relaxations in Mixed Observability Domains
- Optimal trajectories for Bayesian olfactory search in turbulent flows: the low information limit and beyond