activity
20172022
most citedMAVEN: Multi-Agent Variational Exploration

75 citations · 219 across the 9 of their papers we have counts for

collaborators

10 papers

cs.LG20226 cited

Generalization in Cooperative Multi-Agent Systems

Anuj Mahajan, Mikayel Samvelyan, Tarun Gupta +4

Collective intelligence is a fundamental trait shared by several species of living organisms. It has allowed them to thrive in the diverse environmental conditions that exist on ou…

cs.LG20211 cited

Reinforcement Learning in Factored Action Spaces using Tensor Decompositions

Anuj Mahajan, Mikayel Samvelyan, Lei Mao +6

We present an extended abstract for the previously published work TESSERACT [Mahajan et al., 2021], which proposes a novel solution for Reinforcement Learning (RL) in large, factor…

cs.LG20213 cited

Model based Multi-agent Reinforcement Learning with Tensor Decompositions

Pascal Van Der Vaart, Anuj Mahajan, Shimon Whiteson

A challenge in multi-agent reinforcement learning is to be able to generalize over intractable state-action spaces. Inspired from Tesseract [Mahajan et al., 2021], this position pa…

cs.LG202155 cited

Open-Ended Learning Leads to Generally Capable Agents

Open Ended Learning Team, Adam Stooke, Anuj Mahajan +15

In this work we create agents that can perform well beyond a single, individual task, that exhibit much wider generalisation of behaviour to a massive, rich space of challenges. We…

cs.LG20212 cited

SoftDICE for Imitation Learning: Rethinking Off-policy Distribution Matching

Mingfei Sun, Anuj Mahajan, Katja Hofmann +1

We present SoftDICE, which achieves state-of-the-art performance for imitation learning. SoftDICE fixes several key problems in ValueDICE, an off-policy distribution matching appro…

cs.LG20216 cited

Tesseract: Tensorised Actors for Multi-Agent Reinforcement Learning

Anuj Mahajan, Mikayel Samvelyan, Lei Mao +6

Reinforcement Learning in large action spaces is a challenging problem. Cooperative multi-agent reinforcement learning (MARL) exacerbates matters by imposing various constraints on…