activity
20162022
most citedCCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data

246 citations · 448 across the 18 of their papers we have counts for

collaborators

27 papers

cs.CL202217 cited

Consistent Human Evaluation of Machine Translation across Language Pairs

Daniel Licht, Cynthia Gao, Janice Lam +3

Obtaining meaningful quality scores for machine translation systems through human evaluation remains a challenge given the high variability between human evaluators, partly due to…

cs.CL20221 cited

OCR Improves Machine Translation for Low-Resource Languages

Oana Ignat, Jean Maillard, Vishrav Chaudhary +1

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT,…

cs.CL20212 cited

Alternative Input Signals Ease Transfer in Multilingual Machine Translation

Simeng Sun, Angela Fan, James Cross +4

Recent work in multilingual machine translation (MMT) has focused on the potential of positive transfer between languages, particularly cases where higher-resourced languages can b…

cs.CL2021

Classification-based Quality Estimation: Small and Efficient Models for Real-world Applications

Shuo Sun, Ahmed El-Kishky, Vishrav Chaudhary +3

Sentence-level Quality estimation (QE) of machine translation is traditionally formulated as a regression task, and the performance of QE models is typically measured by Pearson co…

cs.CL2021

As Easy as 1, 2, 3: Behavioural Testing of NMT Systems for Numerical Translation

Jun Wang, Chang Xu, Francisco Guzman +3

Mistranslated numbers have the potential to cause serious effects, such as financial loss or medical misinformation. In this work we develop comprehensive assessments of the robust…

cs.CL20211 cited

Putting words into the system's mouth: A targeted attack on neural machine translation using monolingual data poisoning

Jun Wang, Chang Xu, Francisco Guzman +4

Neural machine translation systems are known to be vulnerable to adversarial test inputs, however, as we show in this paper, these systems are also vulnerable to training attacks.…