activity
20192022
collaborators

6 papers

cs.CL2022

From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern French

Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz +4

Language models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historica…

cs.CL2021

Few-shot learning through contextual data augmentation

Farid Arthaud, Rachel Bawden, Alexandra Birch

Machine translation (MT) models used in industries with constantly changing topics, such as translation or news agencies, need to adapt to new data to maintain their performance ov…

cs.CL2020

A Study in Improving BLEU Reference Coverage with Diverse Automatic Paraphrasing

Rachel Bawden, Biao Zhang, Lisa Yankovskaya +2

We investigate a long-perceived shortcoming in the typical use of BLEU: its reliance on a single reference. Using modern neural paraphrasing techniques, we study whether automatica…

cs.CL2019

Document Sub-structure in Neural Machine Translation

Radina Dobreva, Jie Zhou, Rachel Bawden

Current approaches to machine translation (MT) either translate sentences in isolation, disregarding the context they appear in, or model context at the level of the full document,…

cs.CL2019

The University of Edinburgh's Submissions to the WMT19 News Translation Task

Rachel Bawden, Nikolay Bogoychev, Ulrich Germann +4

The University of Edinburgh participated in the WMT19 Shared Task on News Translation in six language directions: English-to-Gujarati, Gujarati-to-English, English-to-Chinese, Chin…

cs.CL2019

DiaBLa: A Corpus of Bilingual Spontaneous Written Dialogues for Machine Translation

Rachel Bawden, Sophie Rosset, Thomas Lavergne +1

We present a new English-French test set for the evaluation of Machine Translation (MT) for informal, written bilingual dialogue. The test set contains 144 spontaneous dialogues (5…