activity
20162024
most citedfairseq: A Fast, Extensible Toolkit for Sequence Modeling

163 citations · 201 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL20222 cited

High-Resource Methodological Bias in Low-Resource Investigations

Maartje ter Hoeve, David Grangier, Natalie Schluter

The central bottleneck for low-resource NLP is typically regarded to be the quantity of accessible data, overlooking the contribution of data quality. This is particularly seen in…

cs.CL20216 cited

On the Complementarity of Data Selection and Fine Tuning for Domain Adaptation

Dan Iter, David Grangier

Domain adaptation of neural networks commonly relies on three training phases: pretraining, selected data training and then fine tuning. Data selection improves target domain gener…

cs.CL2020

Human-Paraphrased References Improve Neural Machine Translation

Markus Freitag, George Foster, David Grangier +1

Automatic evaluation comparing candidate translations to human-generated paraphrases of reference translations has recently been proposed by Freitag et al. When used in place of or…

cs.CL2020

Toward Better Storylines with Sentence-Level Language Models

Daphne Ippolito, David Grangier, Douglas Eck +1

We propose a sentence-level language model which selects the next sentence in a story from a finite set of fluent alternatives. Since it does not need to model fluency, the sentenc…

cs.CL2020

BLEU might be Guilty but References are not Innocent

Markus Freitag, David Grangier, Isaac Caswell

The quality of automatic metrics for machine translation has been increasingly called into question, especially for high-quality systems. This paper demonstrates that, while choice…

cs.CL2019

Translationese as a Language in "Multilingual" NMT

Parker Riley, Isaac Caswell, Markus Freitag +1

Machine translation has an undesirable propensity to produce "translationese" artifacts, which can lead to higher BLEU scores while being liked less by human raters. Motivated by t…