activity
20182024
most citedScaling Language Models: Methods, Analysis & Insights from Training Gopher

243 citations · 507 across the 8 of their papers we have counts for

collaborators

13 papers

cs.CL2024

Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models

Piotr Padlewski, Max Bain, Matthew Henderson +19

We introduce Vibe-Eval: a new open benchmark and framework for evaluating multimodal chat models. Vibe-Eval consists of 269 visual understanding prompts, including 100 of hard diff…

cs.CL2024★ 3 cited

Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reka Team, Aitor Ormazabal, Che Zheng +23

We introduce Reka Core, Flash, and Edge, a series of powerful multimodal language models trained from scratch by Reka. Reka models are able to process and reason with text, images,…

cs.CL2022★ 17 cited

StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models

Adam Liška, Tomáš Kočiský, Elena Gribovskaya +11

Knowledge and language understanding of models evaluated through question answering (QA) has been usually studied on static snapshots of knowledge, like Wikipedia. However, our wor…

cs.CL2022★ 243 cited

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Jack W. Rae, Sebastian Borgeaud, Trevor Cai +77

Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.…

cs.CL2021★ 4 cited

A Systematic Investigation of Commonsense Knowledge in Large Language Models

Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann +3

Language models (LMs) trained on large amounts of data have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup. Here we aim to better understand…

cs.CL2021★ 2 cited

Adaptive Semiparametric Language Models

Dani Yogatama, Cyprien de Masson d'Autume, Lingpeng Kong

We present a language model that combines a large parametric neural network (i.e., a transformer) with a non-parametric episodic memory component in an integrated architecture. Our…