activity
20192025
most citedPost-editing Productivity with Neural Machine Translation: An Empirical Assessment of Speed and Quality in the Banking and Finance Domain

17 citations · 29 across the 8 of their papers we have counts for

collaborators

10 papers

cs.CL2025

A comparison of translation performance between DeepL and Supertext

Alex Flückiger, Chantal Amrhein, Tim Graf +5

As strong machine translation (MT) systems are increasingly based on large language models (LLMs), reliable quality benchmarking requires methods that capture their ability to leve…

cs.CL20242 cited

Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets

Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5

Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…

cs.CL2023

A Benchmark for Evaluating Machine Translation Metrics on Dialects Without Standard Orthography

Noëmi Aepli, Chantal Amrhein, Florian Schottmann +1

For sensible progress in natural language processing, it is important that we are aware of the limitations of the evaluation metrics we use. In this work, we evaluate how robust me…

cs.CL2023

ACES: Translation Accuracy Challenge Sets at WMT 2023

Chantal Amrhein, Nikita Moghe, Liane Guillou

We benchmark the performance of segmentlevel metrics submitted to WMT 2023 using the ACES Challenge Set (Amrhein et al., 2022). The challenge set consists of 36K examples represent…

cs.CL20221 cited

Don't Discard Fixed-Window Audio Segmentation in Speech-to-Text Translation

Chantal Amrhein, Barry Haddow

For real-life applications, it is crucial that end-to-end spoken language translation models perform well on continuous audio, without relying on human-supplied segmentation. For o…

cs.CL20229 cited

ACES: Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics

Chantal Amrhein, Nikita Moghe, Liane Guillou

As machine translation (MT) metrics improve their correlation with human judgement every year, it is crucial to understand the limitations of such metrics at the segment level. Spe…