504 citations · 791 across the 4 of their papers we have counts for
4 papers
Policy Gradient Search: Online Planning and Expert Iteration without Search Trees
Thomas Anthony, Robert Nishihara, Philipp Moritz +2
Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online. During search, the simulation policy is adapted to explore the most promising…
Semi-Supervised Learning by Label Gradient Alignment
Jacob Jackson, John Schulman
We present label gradient alignment, a novel algorithm for semi-supervised learning which imputes labels for the unlabeled data and trains on the imputed labels. We define a semant…
RL: Fast Reinforcement Learning via Slow Reinforcement Learning
Yan Duan, John Schulman, Xi Chen +3
Deep reinforcement learning (deep RL) has been successful in learning sophisticated behaviors automatically; however, the learning process requires a huge number of trials. In cont…
Variational Lossy Autoencoder
Xi Chen, Diederik P. Kingma, Tim Salimans +5
Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good r…