243 citations · 351 across the 5 of their papers we have counts for
6 papers
Solving math word problems with process- and outcome-based feedback
Jonathan Uesato, Nate Kushman, Ramana Kumar +6
Recent work has shown that asking language models to generate reasoning steps improves performance on many reasoning tasks. When moving beyond prompting, this raises the question o…
Teaching language models to support answers with verified quotes
Jacob Menick, Maja Trebacz, Vladimir Mikulik +8
Recent large language models often answer factual questions correctly. But users can't trust any given claim a model makes without fact-checking, because language models can halluc…
Red Teaming Language Models with Language Models
Ethan Perez, Saffron Huang, Francis Song +6
Language Models (LMs) often cannot be deployed because of their potential to harm users in hard-to-predict ways. Prior work identifies harmful behaviors before deployment by using…
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
Synthetic Returns for Long-Term Credit Assignment
David Raposo, Sam Ritter, Adam Santoro +5
Since the earliest days of reinforcement learning, the workhorse method for assigning credit to actions over time has been temporal-difference (TD) learning, which propagates credi…
Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents
Jane X. Wang, Michael King, Nicolas Porcel +14
There has been rapidly growing interest in meta-learning as a method for increasing the flexibility and sample efficiency of reinforcement learning. One problem in this area of res…