1 paper
Marcel Hedman, Kale-ab Abebe Tessera, Juan Claude Formanek +5
Offline multi-agent reinforcement learning (MARL) enables policy learning from fixed datasets, but is prone to coordination failure: agents trained on static, off-policy data conve…