17 citations · 29 across the 8 of their papers we have counts for
10 papers
A comparison of translation performance between DeepL and Supertext
Alex Flückiger, Chantal Amrhein, Tim Graf +5
As strong machine translation (MT) systems are increasingly based on large language models (LLMs), reliable quality benchmarking requires methods that capture their ability to leve…
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5
Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…
A Benchmark for Evaluating Machine Translation Metrics on Dialects Without Standard Orthography
Noëmi Aepli, Chantal Amrhein, Florian Schottmann +1
For sensible progress in natural language processing, it is important that we are aware of the limitations of the evaluation metrics we use. In this work, we evaluate how robust me…
ACES: Translation Accuracy Challenge Sets at WMT 2023
Chantal Amrhein, Nikita Moghe, Liane Guillou
We benchmark the performance of segmentlevel metrics submitted to WMT 2023 using the ACES Challenge Set (Amrhein et al., 2022). The challenge set consists of 36K examples represent…
Don't Discard Fixed-Window Audio Segmentation in Speech-to-Text Translation
Chantal Amrhein, Barry Haddow
For real-life applications, it is crucial that end-to-end spoken language translation models perform well on continuous audio, without relying on human-supplied segmentation. For o…
ACES: Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics
Chantal Amrhein, Nikita Moghe, Liane Guillou
As machine translation (MT) metrics improve their correlation with human judgement every year, it is crucial to understand the limitations of such metrics at the segment level. Spe…