activity
20172022
most citedScaling Language Models: Methods, Analysis & Insights from Training Gopher

243 citations · 597 across the 9 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2022112 cited

Fine-tuning language models to find agreement among humans with diverse preferences

Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan +8

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static…

cs.LG2022133 cited

Improving alignment of dialogue agents via targeted human judgements

Amelia Glaese, Nat McAleese, Maja Trębacz +31

We present Sparrow, an information-seeking dialogue agent trained to be more helpful, correct, and harmless compared to prompted language model baselines. We use reinforcement lear…

cs.LG20209 cited

Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning

Giambattista Parascandolo, Lars Buesing, Josh Merel +6

Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumpt…

cs.LG2019

Behaviour Suite for Reinforcement Learning

Ian Osband, Yotam Doron, Matteo Hessel +11

This paper introduces the Behaviour Suite for Reinforcement Learning, or bsuite for short. bsuite is a collection of carefully-designed experiments that investigate core capabiliti…

cs.LG2019

When to use parametric models in reinforcement learning?

Hado van Hasselt, Matteo Hessel, John Aslanides

We examine the question of when and how parametric models are most useful in reinforcement learning. In particular, we look at commonalities and differences between parametric mode…

cs.LG201921 cited

TF-Replicator: Distributed Machine Learning for Researchers

Peter Buchlovsky, David Budden, Dominik Grewe +9

We describe TF-Replicator, a framework for distributed machine learning designed for DeepMind researchers and implemented as an abstraction over TensorFlow. TF-Replicator simplifie…