4 papers
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
Asim Osman, Sasha Abramowitz, Mark Bergh +13
Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-cra…
CODA: Coordination via On-Policy Diffusion for Multi-Agent Offline Reinforcement Learning
Marcel Hedman, Kale-ab Abebe Tessera, Juan Claude Formanek +5
Offline multi-agent reinforcement learning (MARL) enables policy learning from fixed datasets, but is prone to coordination failure: agents trained on static, off-policy data conve…
Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARL
Claude Formanek, Omayma Mahjoub, Louay Ben Nessir +10
A key challenge in offline multi-agent reinforcement learning (MARL) is achieving effective many-agent multi-step coordination in complex environments. In this work, we propose Ory…
Sable: a Performant, Efficient and Scalable Sequence Model for MARL
Omayma Mahjoub, Sasha Abramowitz, Ruan de Kock +8
As multi-agent reinforcement learning (MARL) progresses towards solving larger and more complex problems, it becomes increasingly important that algorithms exhibit the key properti…