most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 101 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across Paragraphs

Juan Diego Rodriguez, Katrin Erk, Greg Durrett

Understanding when two pieces of text convey the same information is a goal touching many subproblems in NLP, including textual entailment and fact-checking. This problem becomes m…

cs.CL2023★ 1 cited

WiCE: Real-World Entailment for Claims in Wikipedia

Ryo Kamoi, Tanya Goyal, Juan Diego Rodriguez +1

Textual entailment models are increasingly applied in settings like fact-checking, presupposition verification in question answering, or summary evaluation. However, these represen…

cs.CL2021★ 25 cited

NL-Augmenter: A Framework for Task-Sensitive Natural Language Augmentation

Kaustubh D. Dhole, Varun Gangal, Sebastian Gehrmann +122

Data augmentation is an important component in the robustness evaluation of models in natural language processing (NLP) and in enhancing the diversity of the data they are trained…

cs.DB2021★ 23 cited

Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards

Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez +3

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of…

cs.CL2021★ 52 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…