45 citations · 80 across the 3 of their papers we have counts for
3 papers
cs.LG2023★ 31 cited
Let's Verify Step by Step
Hunter Lightman, Vineet Kosaraju, Yura Burda +7
In recent years, large language models have greatly improved in their ability to perform complex multi-step reasoning. However, even state-of-the-art models still regularly produce…
cs.LG2023★ 4 cited
Scaling laws for single-agent reinforcement learning
Jacob Hilton, Jie Tang, John Schulman
Recent work has shown that, in generative modeling, cross-entropy loss improves smoothly with model size and training compute, following a power law plus constant scaling law. One…
cs.CL2022★ 45 cited
Efficient Training of Language Models to Fill in the Middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak +4
We show that autoregressive language models can learn to infill text after we apply a straightforward transformation to the dataset, which simply moves a span of text from the midd…