activity
20162022
most citedThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

52 citations · 69 across the 7 of their papers we have counts for

collaborators

12 papers

cs.CL20222 cited

PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation

Sedrick Scott Keh, Kevin Lu, Varun Gangal +4

A personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of person…

cs.CL202112 cited

Automatic Construction of Evaluation Suites for Natural Language Generation Datasets

Simon Mille, Kaustubh D. Dhole, Saad Mahamood +5

Machine learning approaches applied to NLP are often evaluated by summarizing their performance in a single number, for example accuracy. Since most test sets are constructed as an…

cs.CL2021

Improving Automated Evaluation of Open Domain Dialog via Diverse Reference Augmentation

Varun Gangal, Harsh Jhamtani, Eduard Hovy +1

Multiple different responses are often plausible for a given open domain dialog context. Prior work has shown the importance of having multiple valid reference responses for meanin…

cs.CL202152 cited

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal +53

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of auto…

cs.CL2020

GenAug: Data Augmentation for Finetuning Text Generators

Steven Y. Feng, Varun Gangal, Dongyeop Kang +2

In this paper, we investigate data augmentation for text generation, which we call GenAug. Text generation and language modeling are important tasks within natural language process…

cs.CL2020

BERTering RAMS: What and How Much does BERT Already Know About Event Arguments? -- A Study on the RAMS Dataset

Varun Gangal, Eduard Hovy

Using the attention map based probing frame-work from (Clark et al., 2019), we observe that, on the RAMS dataset (Ebner et al., 2020), BERT's attention heads have modest but well a…