7 citations · 24 across the 18 of their papers we have counts for
32 papers · 1 filter
IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages
Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare +14
Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and lang…
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data
Haoran Sun, Renren Jin, Shaoyang Xu +10
Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
Matthias Sperber, Ondřej Bojar, Barry Haddow +8
Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists o…
Who Are We Talking About? Handling Person Names in Speech Translation
Marco Gaido, Matteo Negri, Marco Turchi
Recent work has shown that systems for speech translation (ST) -- similarly to automatic speech recognition (ASR) -- poorly handle person names. This shortcoming does not only lead…
Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli +2
Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages.…
Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation
Marco Gaido, Susana Rodríguez, Matteo Negri +2
Automatic translation systems are known to struggle with rare words. Among these, named entities (NEs) and domain-specific terms are crucial, since errors in their translation can…