3 papers
cs.LG2026
Utility-Constrained Policy Optimization
Mehrdad Moghimi, Bernardo Avila Pires
Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents; however, the framework does not support risk-sensitive constraints. This can be pro…
cs.LG2025
Optimizing Return Distributions with Distributional Dynamic Programming
Bernardo Ãvila Pires, Mark Rowland, Diana Borsa +6
We introduce distributional dynamic programming (DP) methods for optimizing statistical functionals of the return distribution, with standard reinforcement learning as a special ca…
cs.LG2025
Representation Learning via Non-Contrastive Mutual Information
Zhaohan Daniel Guo, Bernardo Avila Pires, Khimya Khetarpal +2
Labeling data is often very time consuming and expensive, leaving us with a majority of unlabeled data. Self-supervised representation learning methods such as SimCLR (Chen et al.,…