Incremental Clustering and Expansion for Faster Optimal Planning in Dec-POMDPs
arXiv:1402.0566 · doi:10.1613/jair.3804
Abstract
This article presents the state-of-the-art in optimal solution methods for decentralized partially observable Markov decision processes (Dec-POMDPs), which are general models for collaborative multiagent planning under uncertainty. Building off the generalized multiagent A* (GMAA*) algorithm, which reduces the problem to a tree of one-shot collaborative Bayesian games (CBGs), we describe several advances that greatly expand the range of Dec-POMDPs that can be solved optimally. First, we introduce lossless incremental clustering of the CBGs solved by GMAA*, which achieves exponential speedups without sacrificing optimality. Second, we introduce incremental expansion of nodes in the GMAA* search tree, which avoids the need to expand all children, the number of which is in the worst case doubly exponential in the nodes depth. This is particularly beneficial when little clustering is possible. In addition, we introduce new hybrid heuristic representations that are more compact and thereby enable the solution of larger Dec-POMDPs. We provide theoretical guarantees that, when a suitable heuristic is used, both incremental clustering and incremental expansion yield algorithms that are both complete and search equivalent. Finally, we present extensive empirical results demonstrating that GMAA*-ICE, an algorithm that synthesizes these advances, can optimally solve Dec-POMDPs of unprecedented size.
References in corpus (13)
- Value-Function Approximations for Partially Observable Markov Decision Processes
- Online Planning Algorithms for POMDPs
- The Communicative Multiagent Team Decision Problem: Analyzing Teamwork Theories and Models
- Incremental Pruning: A Simple, Fast, Exact Method for Partially Observable Markov Decision Processes
- Decentralized Control of Cooperative Systems: Categorization and Complexity Analysis
- MAA*: A Heuristic Search Algorithm for Solving Decentralized POMDPs
- Improved Memory-Bounded Dynamic Programming for Decentralized POMDPs
- Policy Iteration for Decentralized Control of Markov Decision Processes
- Monte Carlo Sampling Methods for Approximating Interactive POMDPs
- Optimizing Memory-Bounded Controllers for Decentralized POMDPs
- Anytime Planning for Decentralized POMDPs using Expectation Maximization
- An Investigation into Mathematical Programming for Finite Horizon Decentralized POMDPs
- Rollout Sampling Policy Iteration for Decentralized POMDPs