65 citations · 212 across the 16 of their papers we have counts for
24 papers
You May Not Need Ratio Clipping in PPO
Mingfei Sun, Vitaly Kurin, Guoqing Liu +4
Proximal Policy Optimization (PPO) methods learn a policy by iteratively performing multiple mini-batch optimization epochs of a surrogate objective with one set of sampled data. R…
NeurIPS 2021 Competition IGLU: Interactive Grounded Language Understanding in a Collaborative Environment
Julia Kiseleva, Ziming Li, Mohammad Aliannejadi +12
Human intelligence has the remarkable ability to adapt to new tasks and environments quickly. Starting from a very young age, humans acquire new skills and learn how to solve new t…
SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents
Grgur Kovač, Rémy Portelas, Katja Hofmann +1
Building embodied autonomous agents capable of participating in social interactions with humans is one of the main challenges in AI. Within the Deep Reinforcement Learning (DRL) fi…
Strategically Efficient Exploration in Competitive Multi-agent Reinforcement Learning
Robert Loftin, Aadirupa Saha, Sam Devlin +1
High sample complexity remains a barrier to the application of reinforcement learning (RL), particularly in multi-agent systems. A large body of work has demonstrated that explorat…
SoftDICE for Imitation Learning: Rethinking Off-policy Distribution Matching
Mingfei Sun, Anuj Mahajan, Katja Hofmann +1
We present SoftDICE, which achieves state-of-the-art performance for imitation learning. SoftDICE fixes several key problems in ValueDICE, an off-policy distribution matching appro…
Grounding Spatio-Temporal Language with Transformers
Tristan Karch, Laetitia Teodorescu, Katja Hofmann +2
Language is an interface to the outside world. In order for embodied agents to use it, language must be grounded in other, sensorimotor modalities. While there is an extended liter…