52 citations · 53 across the 2 of their papers we have counts for
3 papers
Underreporting of errors in NLG output, and what to do about it
Emiel van Miltenburg, Miruna-Adriana Clinciu, Ondřej Dušek +8
We observe a severe under-reporting of the different kinds of errors that Natural Language Generation systems make. This is a problem, because mistakes are an important indicator o…
The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53
We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…
A Study of Automatic Metrics for the Evaluation of Natural Language Explanations
Miruna Clinciu, Arash Eshghi, Helen Hastie
As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural lang…