9 citations · 11 across the 6 of their papers we have counts for
10 papers
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Laurie Burchell, Ona de Gibert, Nikolay Arefyev +32
Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In th…
Pitfalls and Outlooks in Using COMET
Vilém Zouhar, Pinzhen Chen, Tsz Kin Lam +2
The COMET metric has blazed a trail in the machine translation community, given its strong correlation with human judgements of translation quality. Its success stems from being a…
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5
Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…
ACES: Translation Accuracy Challenge Sets at WMT 2023
Chantal Amrhein, Nikita Moghe, Liane Guillou
We benchmark the performance of segmentlevel metrics submitted to WMT 2023 using the ACES Challenge Set (Amrhein et al., 2022). The challenge set consists of 36K examples represent…
Interpreting User Requests in the Context of Natural Language Standing Instructions
Nikita Moghe, Patrick Xia, Jacob Andreas +3
Users of natural language interfaces, generally powered by Large Language Models (LLMs),often must repeat their preferences each time they make a similar request. We describe an ap…
ACES: Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics
Chantal Amrhein, Nikita Moghe, Liane Guillou
As machine translation (MT) metrics improve their correlation with human judgement every year, it is crucial to understand the limitations of such metrics at the segment level. Spe…