activity
20172022
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 65 across the 6 of their papers we have counts for

collaborators

12 papers

cs.CL20222 cited

Dialect-robust Evaluation of Generated Text

Jiao Sun, Thibault Sellam, Elizabeth Clark +6

Evaluation metrics that are not robust to dialect variation make it impossible to tell how well systems perform for many groups of users, and can even penalize systems for producin…

cs.CL20228 cited

Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text

Sebastian Gehrmann, Elizabeth Clark, Thibault Sellam

Evaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted. This issue has become more urgent, si…

cs.CL2021

Learning Compact Metrics for MT

Amy Pu, Hyung Won Chung, Ankur P. Parikh +2

Recent developments in machine translation and multilingual text generation have led researchers to adopt trained metrics such as COMET or BLEURT, which treat evaluation as a regre…

cs.CL202152 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…

cs.CL2020

Learning to Evaluate Translation Beyond English: BLEURT Submissions to the WMT Metrics 2020 Shared Task

Thibault Sellam, Amy Pu, Hyung Won Chung +5

The quality of machine translation systems has dramatically improved over the last decade, and as a result, evaluation has become an increasingly challenging problem. This paper de…

cs.CL2020

BLEURT: Learning Robust Metrics for Text Generation

Thibault Sellam, Dipanjan Das, Ankur P. Parikh

Text generation has made significant advances in the last few years. Yet, evaluation metrics have lagged behind, as the most popular choices (e.g., BLEU and ROUGE) may correlate po…