8 papers
Risk-Aware General-Utility Markov Decision Processes
Pedro P. Santos, Fábio Vital, Alberto Sardinha +1
We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective…
Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning
Pedro P. Santos, Alberto Sardinha, Francisco S. Melo
In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the ag…
Entropic Risk-Aware Monte Carlo Tree Search
Pedro P. Santos, Jacopo Silvestrin, Alberto Sardinha +1
We propose a provably correct Monte Carlo tree search (MCTS) algorithm for solving risk-aware Markov decision processes (MDPs) with entropic risk measure (ERM) objectives. We provi…
The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes
Pedro P. Santos, Alberto Sardinha, Francisco S. Melo
The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of sta…
RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains
João G. Ribeiro, Yaniv Oren, Alberto Sardinha +2
This paper proposes RecBayes, a novel approach for ad hoc teamwork under partial observability, a setting where agents are deployed on-the-fly to environments where pre-existing te…
Implicit Repair with Reinforcement Learning in Emergent Communication
Fábio Vital, Alberto Sardinha, Francisco S. Melo
Conversational repair is a mechanism used to detect and resolve miscommunication and misinformation problems when two or more agents interact. One particular and underexplored form…