activity
20172022
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 257 across the 15 of their papers we have counts for

collaborators

23 papers

cs.CL2021

LMdiff: A Visual Diff Tool to Compare Language Models

Hendrik Strobelt, Benjamin Hoover, Arvind Satyanarayan +1

While different language models are ubiquitous in NLP, it is hard to contrast their outputs and identify which contexts one can handle better than the other. To address this questi…

cs.CL2021

Learning Compact Metrics for MT

Amy Pu, Hyung Won Chung, Ankur P. Parikh +2

Recent developments in machine translation and multilingual text generation have led researchers to adopt trained metrics such as COMET or BLEURT, which treat evaluation as a regre…

cs.DB202123 cited

Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards

Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez +3

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of…

cs.CL20211 cited

Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models

Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann +3

Targeted syntactic evaluations have demonstrated the ability of language models to perform subject-verb agreement given difficult contexts. To elucidate the mechanisms by which the…

cs.CL202112 cited

Automatic Construction of Evaluation Suites for Natural Language Generation Datasets

Simon Mille, Kaustubh D. Dhole, Saad Mahamood +5

Machine learning approaches applied to NLP are often evaluated by summarizing their performance in a single number, for example accuracy. Since most test sets are constructed as an…

cs.CL202152 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…