activity
20162026
most citedMastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

1.1k citations · 3.3k across the 17 of their papers we have counts for

collaborators
Showing 2022Show all

6 papers · 1 filter

cs.CL2022★ 21 cited

Self-conditioned Embedding Diffusion for Text Generation

Robin Strudel, Corentin Tallec, Florent Altché +8

Can continuous diffusion models bring the same performance breakthrough on natural language they did for image generation? To circumvent the discrete nature of text data, we can si…

cs.AI2022★ 153 cited

Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning

Julien Perolat, Bart de Vylder, Daniel Hennes +31

We introduce DeepNash, an autonomous agent capable of learning to play the imperfect information game Stratego from scratch, up to a human expert level. Stratego is one of the few…

cs.LG2022★ 11 cited

Large-Scale Retrieval for Reinforcement Learning

Peter C. Humphreys, Arthur Guez, Olivier Tieleman +3

Effective decision making involves flexibly relating past experiences and relevant contextual information to a novel situation. In deep reinforcement learning (RL), the dominant pa…

cs.CL2022★ 672 cited

Training Compute-Optimal Large Language Models

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch +19

We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget. We find that current large language models are si…

cs.CL2022★ 23 cited

Unified Scaling Laws for Routed Language Models

Aidan Clark, Diego de las Casas, Aurelia Guy +23

The performance of a language model has been shown to be effectively modeled as a power-law in its parameter count. Here we study the scaling behaviors of Routing Networks: archite…

cs.CL2022★ 243 cited

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…