Online Planning Algorithms for POMDPs
arXiv:1401.3436 · doi:10.1613/jair.2567
Abstract
Partially Observable Markov Decision Processes (POMDPs) provide a rich framework for sequential decision-making under uncertainty in stochastic domains. However, solving a POMDP is often intractable except for small problems due to their complexity. Here, we focus on online approaches that alleviate the computational complexity by computing good local policies at each decision step during the execution. Online algorithms generally consist of a lookahead search to find the best action to execute at each time step in an environment. Our objectives here are to survey the various existing online POMDP methods, analyze their properties and discuss their advantages and disadvantages; and to thoroughly evaluate these online approaches in different environments under various metrics (return, error bound reduction, lower bound improvement). Our experimental results indicate that state-of-the-art online heuristic search methods can handle large POMDP domains efficiently.
References in corpus (8)
- Value-Function Approximations for Partially Observable Markov Decision Processes
- Online Planning Algorithms for POMDPs
- Tractable Inference for Complex Stochastic Processes
- Incremental Pruning: A Simple, Fast, Exact Method for Partially Observable Markov Decision Processes
- Solving POMDPs by Searching in Policy Space
- Speeding Up the Convergence of Value Iteration in Partially Observable Markov Decision Processes
- Point-Based POMDP Algorithms: Improved Analysis and Implementation
- Approximate Planning for Factored POMDPs using Belief State Simplification
Cited by in corpus (52)
- Online Planning Algorithms for POMDPs
- DESPOT: Online POMDP Planning with Regularization
- Partially Observable Markov Decision Processes in Robotics: A Survey
- Partially Observable Markov Decision Processes (POMDPs) and Robotics
- Planning for robotic exploration based on forward simulation
- Efficient Planning under Uncertainty with Macro-actions
- Monte Carlo Sampling Methods for Approximating Interactive POMDPs
- Searching for a source without gradients: how good is infotaxis and how to beat it
- Incremental Clustering and Expansion for Faster Optimal Planning in Dec-POMDPs
- Human-Centric Resource Allocation for the Metaverse With Multiaccess Edge Computing
- Exploiting Submodular Value Functions For Scaling Up Active Perception
- Hilbert Space Embeddings of POMDPs
- Stochastic Motion Planning under Partial Observability for Mobile Robots with Continuous Range Measurements
- Learning Belief Representations for Imitation Learning in POMDPs
- pomdp_py: A Framework to Build and Solve POMDP Problems
- Continuous Relaxation of Symbolic Planner for One-Shot Imitation Learning
- Optimality Guarantees for Particle Belief Approximation of POMDPs
- FHHOP: A Factored Hybrid Heuristic Online Planning Algorithm for Large POMDPs
- Learning Deep Neural Network Policies with Continuous Memory States
- Active Inference Tree Search in Large POMDPs
- A Model-Based, Decision-Theoretic Perspective on Automated Cyber Response
- Efficient Uncertainty-aware Decision-making for Automated Driving Using Guided Branching
- Optimal Sensing via Multi-armed Bandit Relaxations in Mixed Observability Domains
- Hindsight is Only 50/50: Unsuitability of MDP based Approximate POMDP Solvers for Multi-resolution Information Gathering
- Using Social Networks to Aid Homeless Shelters: Dynamic Influence Maximization under Uncertainty - An Extended Version
- Planning to Chronicle
- Model-based Bayesian Reinforcement Learning for Dialogue Management
- POMDP-lite for Robust Robot Planning under Uncertainty
- Online Learning and Planning in Partially Observable Domains without Prior Knowledge
- Scalable Planning and Learning for Multiagent POMDPs: Extended Version
- Learning in POMDPs with Monte Carlo Tree Search
- Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
- Myopic Policy Bounds for Information Acquisition POMDPs
- Sequential Bayesian Optimisation as a POMDP for Environment Monitoring with UAVs
- Cooperative Trajectory Planning in Uncertain Environments with Monte Carlo Tree Search and Risk Metrics
- On State Variables, Bandit Problems and POMDPs
- Stochastic 2-D Motion Planning with a POMDP Framework
- Grasping and Manipulation with a Multi-Fingered Hand
- Robotic manipulation of multiple objects as a POMDP
- Combining Offline Models and Online Monte-Carlo Tree Search for Planning from Scratch
- Batch Belief Trees for Motion Planning Under Uncertainty
- Reputation-driven Decision-making in Networks of Stochastic Agents
- Planning with Submodular Objective Functions
- Deep Reinforcement Learning for High Level Character Control
- Minimum-Latency FEC Design with Delayed Feedback: Mathematical Modeling and Efficient Algorithms
- Social Navigation Planning Based on People's Awareness of Robots
- Rollout Heuristics for Online Stochastic Contingent Planning
- A Bayesian Approach to Identifying Representational Errors
- Efficient Sampling-Based Maximum Entropy Inverse Reinforcement Learning with Application to Autonomous Driving
- Deceptive Kernel Function on Observations of Discrete POMDP
- Active Goal Recognition
- Memory Bounded Open-Loop Planning in Large POMDPs using Thompson Sampling