activity
20192023
most citedVideo PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

50 citations · 84 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2022★ 50 cited

Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos

Bowen Baker, Ilge Akkaya, Peter Zhokhov +6

Pretraining on noisy, internet-scale datasets has been heavily studied as a technique for training models with broad, general capabilities for text, images, and other modalities. H…

cs.LG2021★ 4 cited

The MineRL BASALT Competition on Learning from Human Feedback

Rohin Shah, Cody Wild, Steven H. Wang +10

The last decade has seen a significant increase of interest in deep learning research, with many public successes that have demonstrated its potential. As such, these systems are n…

cs.LG2021★ 9 cited

Multi-task curriculum learning in a complex, visual, hard-exploration domain: Minecraft

Ingmar Kanitscheider, Joost Huizinga, David Farhi +9

An important challenge in reinforcement learning is training agents that can solve a wide variety of tasks. If tasks depend on each other (e.g. needing to learn to walk before lear…

cs.LG2021★ 1 cited

Towards robust and domain agnostic reinforcement learning competitions

William Hebgen Guss, Stephanie Milani, Nicholay Topin +26

Reinforcement learning competitions have formed the basis for standard research benchmarks, galvanized advances in the state-of-the-art, and shaped the direction of the field. Desp…

cs.LG2021★ 14 cited

The MineRL 2020 Competition on Sample Efficient Reinforcement Learning using Human Priors

William H. Guss, Mario Ynocente Castro, Sam Devlin +12

Although deep reinforcement learning has led to breakthroughs in many difficult domains, these successes have required an ever-increasing number of samples, affording only a shrink…

cs.LG2020★ 4 cited

Guaranteeing Reproducibility in Deep Learning Competitions

Brandon Houghton, Stephanie Milani, Nicholay Topin +5

To encourage the development of methods with reproducible and robust training behavior, we propose a challenge paradigm where competitors are evaluated directly on the performance…