1 paper
Francesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi +1
We study online learning in \emph{constrained MDPs} (CMDPs), focusing on the goal of attaining sublinear strong regret and strong cumulative constraint violation. Differently from…