2 papers
cs.LG2026
How Log-Barrier Helps Exploration in Policy Optimization
Leonardo Cesani, Matteo Papini, Marcello Restelli
Recently, it has been shown that the Stochastic Gradient Bandit (SGB) algorithm converges to a globally optimal policy with a constant learning rate. However, these guarantees rely…
cs.LG2025
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes
Alessandro Montenegro, Leonardo Cesani, Marco Mussi +2
Constrained Reinforcement Learning (CRL) addresses sequential decision-making problems where agents are required to achieve goals by maximizing the expected return while meeting do…