activity
20182025
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 901 across the 18 of their papers we have counts for

collaborators
Showing 2021 · cs.CLShow all

6 papers · 2 filters

cs.CL2021

Improving Compositional Generalization with Self-Training for Data-to-Text Generation

Sanket Vaibhav Mehta, Jinfeng Rao, Yi Tay +3

Data-to-text generation focuses on generating fluent natural language responses from structured meaning representations (MRs). Such representations are compositional and it is cost…

cs.CL2021

Using Machine Translation to Localize Task Oriented NLG Output

Scott Roy, Cliff Brunk, Kyu-Young Kim +6

One of the challenges in a task oriented natural language application like the Google Assistant, Siri, or Alexa is to localize the output to many languages. This paper explores doi…

cs.CL2021★ 12 cited

Automatic Construction of Evaluation Suites for Natural Language Generation Datasets

Simon Mille, Kaustubh D. Dhole, Saad Mahamood +5

Machine learning approaches applied to NLP are often evaluated by summarizing their performance in a single number, for example accuracy. Since most test sets are constructed as an…

cs.CL2021★ 1 cited

nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?

Mihir Kale, Aditya Siddhant, Noah Constant +3

Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In…

cs.CL2021

ByT5: Towards a token-free future with pre-trained byte-to-byte models

Linting Xue, Aditya Barua, Noah Constant +5

Most widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw te…

cs.CL2021★ 52 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…