13.4k citations · 18.5k across the 12 of their papers we have counts for
5 papers · 1 filter
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun +55
We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex p…
Variational Lossy Autoencoder
Xi Chen, Diederik P. Kingma, Tim Salimans +5
Representation learning seeks to expose certain aspects of observed data in a learned representation that's amenable to downstream tasks like classification. For instance, a good r…
Learning Online Alignments with Continuous Rewards Policy Gradient
Yuping Luo, Chung-Cheng Chiu, Navdeep Jaitly +1
Sequence-to-sequence models with soft attention had significant success in machine translation, speech recognition, and question answering. Though capable and easy to use, they req…
Move Evaluation in Go Using Deep Convolutional Neural Networks
Chris J. Maddison, Aja Huang, Ilya Sutskever +1
The game of Go is more challenging than other board games, due to the difficulty of constructing a position or move evaluation function. In this paper we investigate whether deep c…
Estimating the Hessian by Back-propagating Curvature
James Martens, Ilya Sutskever, Kevin Swersky
In this work we develop Curvature Propagation (CP), a general technique for efficiently computing unbiased approximations of the Hessian of any function that is computed using a co…