collaborators

8 papers

cs.LG2026

Risk-Aware General-Utility Markov Decision Processes

Pedro P. Santos, Fábio Vital, Alberto Sardinha +1

We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk measure of the distribution of objective…

cs.LG2026

Solving General-Utility Markov Decision Processes in the Single-Trial Regime with Online Planning

Pedro P. Santos, Alberto Sardinha, Francisco S. Melo

In this work, we contribute the first approach to solve infinite-horizon discounted general-utility Markov decision processes (GUMDPs) in the single-trial regime, i.e., when the ag…

cs.LG2026

Entropic Risk-Aware Monte Carlo Tree Search

Pedro P. Santos, Jacopo Silvestrin, Alberto Sardinha +1

We propose a provably correct Monte Carlo tree search (MCTS) algorithm for solving risk-aware Markov decision processes (MDPs) with entropic risk measure (ERM) objectives. We provi…

cs.LG2025

The Number of Trials Matters in Infinite-Horizon General-Utility Markov Decision Processes

Pedro P. Santos, Alberto Sardinha, Francisco S. Melo

The general-utility Markov decision processes (GUMDPs) framework generalizes the MDPs framework by considering objective functions that depend on the frequency of visitation of sta…

cs.MA2025

RecBayes: Recurrent Bayesian Ad Hoc Teamwork in Large Partially Observable Domains

João G. Ribeiro, Yaniv Oren, Alberto Sardinha +2

This paper proposes RecBayes, a novel approach for ad hoc teamwork under partial observability, a setting where agents are deployed on-the-fly to environments where pre-existing te…

cs.LG2025

Implicit Repair with Reinforcement Learning in Emergent Communication

Fábio Vital, Alberto Sardinha, Francisco S. Melo

Conversational repair is a mechanism used to detect and resolve miscommunication and misinformation problems when two or more agents interact. One particular and underexplored form…