6 papers
From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz +4
Language models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historica…
Few-shot learning through contextual data augmentation
Farid Arthaud, Rachel Bawden, Alexandra Birch
Machine translation (MT) models used in industries with constantly changing topics, such as translation or news agencies, need to adapt to new data to maintain their performance ov…
A Study in Improving BLEU Reference Coverage with Diverse Automatic Paraphrasing
Rachel Bawden, Biao Zhang, Lisa Yankovskaya +2
We investigate a long-perceived shortcoming in the typical use of BLEU: its reliance on a single reference. Using modern neural paraphrasing techniques, we study whether automatica…
Document Sub-structure in Neural Machine Translation
Radina Dobreva, Jie Zhou, Rachel Bawden
Current approaches to machine translation (MT) either translate sentences in isolation, disregarding the context they appear in, or model context at the level of the full document,…
The University of Edinburgh's Submissions to the WMT19 News Translation Task
Rachel Bawden, Nikolay Bogoychev, Ulrich Germann +4
The University of Edinburgh participated in the WMT19 Shared Task on News Translation in six language directions: English-to-Gujarati, Gujarati-to-English, English-to-Chinese, Chin…
DiaBLa: A Corpus of Bilingual Spontaneous Written Dialogues for Machine Translation
Rachel Bawden, Sophie Rosset, Thomas Lavergne +1
We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5…