4.3k citations · 4.9k across the 11 of their papers we have counts for
3 papers · 1 filter
Leveraging Procedural Generation to Benchmark Reinforcement Learning
Karl Cobbe, Christopher Hesse, Jacob Hilton +1
We introduce Procgen Benchmark, a suite of 16 procedurally generated game-like environments designed to benchmark both sample efficiency and generalization in reinforcement learnin…
Policy Gradient Search: Online Planning and Expert Iteration without Search Trees
Thomas Anthony, Robert Nishihara, Philipp Moritz +2
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online. During search, the simulation policy is adapted to explore the most promising…
Semi-Supervised Learning by Label Gradient Alignment
Jacob Jackson, John Schulman
We present label gradient alignment, a novel algorithm for semi-supervised learning which imputes labels for the unlabeled data and trains on the imputed labels. We define a semant…