31 citations · 45 across the 3 of their papers we have counts for
7 papers · 1 filter
Let's Verify Step by Step
Hunter Lightman, Vineet Kosaraju, Yura Burda +7
In recent years, large language models have greatly improved in their ability to perform complex multi-step reasoning. However, even state-of-the-art models still regularly produce…
Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft
Ingmar Kanitscheider, Joost Huizinga, David Farhi +9
An important challenge in reinforcement learning is training agents that can solve a wide variety of tasks. If tasks depend on each other (e.g. needing to learn to walk before lear…
Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences
Bowen Baker
Multi-agent reinforcement learning (MARL) has shown recent success in increasingly complex fixed-team zero-sum environments. However, the real world is not zero-sum nor does it hav…
Emergent Tool Use From Multi-Agent Autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov +4
Through multi-agent competition, the simple objective of hide-and-seek, and standard reinforcement learning algorithms at scale, we find that agents create a self-supervised autocu…
Learning Dexterous In-Hand Manipulation
OpenAI, Marcin Andrychowicz, Bowen Baker +14
We use reinforcement learning (RL) to learn dexterous in-hand manipulation policies which can perform vision-based object reorientation on a physical Shadow Dexterous Hand. The tra…
Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
Matthias Plappert, Marcin Andrychowicz, Alex Ray +9
The purpose of this technical report is two-fold. First of all, it introduces a suite of challenging continuous control tasks (integrated with OpenAI Gym) based on currently existi…