2 papers
cs.LG2026
Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
Oussama Hidaoui, Omer Ebead, Ulrich Armel Mbou Sob +14
Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task…
cs.LG2026
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
Asim Osman, Sasha Abramowitz, Mark Bergh +13
Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-cra…