243 citations · 254 across the 5 of their papers we have counts for
7 papers
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77
Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…
You should evaluate your language model on marginal likelihood over tokenisations
Kris Cao, Laura Rimell
Neural language models typically tokenise input text into sub-word units to achieve an open vocabulary. The standard approach is to use a single canonical tokenisation at both trai…
Pretraining the Noisy Channel Model for Task-Oriented Dialogue
Qi Liu, Lei Yu, Laura Rimell +1
Direct decoding for task-oriented dialogue is known to suffer from the explaining-away effect, manifested in models that prefer short and generic responses. Here we argue for the u…
Probing Emergent Semantics in Predictive Agents via Question Answering
Abhishek Das, Federico Carnevale, Hamza Merzic +8
Recent work has shown how predictive modeling can endow agents with rich knowledge of their surroundings, improving their ability to act in complex environments. We propose questio…
Syntactic Structure Distillation Pretraining For Bidirectional Encoders
Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried +4
Textual representation learners trained on large amounts of data have achieved notable success on downstream tasks; intriguingly, they have also performed well on challenging tests…
Neural Generative Rhetorical Structure Parsing
Amandla Mabona, Laura Rimell, Stephen Clark +1
Rhetorical structure trees have been shown to be useful for several document-level tasks including summarization and document classification. Previous approaches to RST parsing hav…