3 papers
cs.MA2026
Risk-seeking conservative policy iteration with agent-state based policies for Dec-POMDPs with guaranteed convergence
Amit Sinha, Matthieu Geist, Aditya Mahajan
Optimally solving decentralized decision-making problems modeled as Dec-POMDPs is known to be NEXP-complete. These optimal solutions are policies based on the entire history of obs…
cs.LG2025
Convergence of regularized agent-state-based Q-learning in POMDPs
Amit Sinha, Matthieu Geist, Aditya Mahajan
In this paper, we present a framework to understand the convergence of commonly used Q-learning reinforcement learning algorithms in practice. Two salient features of such algorith…
cs.LG2024
Periodic agent-state based Q-learning for POMDPs
Amit Sinha, Matthieu Geist, Aditya Mahajan
The standard approach for Partially Observable Markov Decision Processes (POMDPs) is to convert them to a fully observed belief-state MDP. However, the belief state depends on the…