185 citations · 262 across the 8 of their papers we have counts for
4 papers · 1 filter
UniMASK: Unified Inference in Sequential Decision Problems
Micah Carroll, Orr Paradise, Jessy Lin +8
Randomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same…
Generalization in Cooperative Multi-Agent Systems
Anuj Mahajan, Mikayel Samvelyan, Tarun Gupta +4
Collective intelligence is a fundamental trait shared by several species of living organisms. It has allowed them to thrive in the diverse environmental conditions that exist on ou…
You May Not Need Ratio Clipping in PPO
Mingfei Sun, Vitaly Kurin, Guoqing Liu +4
Proximal Policy Optimization (PPO) methods learn a policy by iteratively performing multiple mini-batch optimization epochs of a surrogate objective with one set of sampled data. R…
Adversarial Imitation Learning from Incomplete Demonstrations
Mingfei Sun, Xiaojuan Ma
Imitation learning targets deriving a mapping from states to actions, a.k.a. policy, from expert demonstrations. Existing methods for imitation learning typically require any actio…