activity
20172026
most citedAdapting Multilingual Neural Machine Translation to Unseen Languages

7 citations · 24 across the 18 of their papers we have counts for

collaborators
Showing cs.CLShow all

32 papers · 1 filter

cs.CL2026

IndicQE-APE: A Consolidated Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare +14

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and lang…

cs.CL2024

FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data

Haoran Sun, Renren Jin, Shaoyang Xu +10

Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…

cs.CL2024

Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation

Matthias Sperber, Ondřej Bojar, Barry Haddow +8

Human evaluation is a critical component in machine translation system development and has received much attention in text translation research. However, little prior work exists o…

cs.CL2022

Who Are We Talking About? Handling Person Names in Speech Translation

Marco Gaido, Matteo Negri, Marco Turchi

Recent work has shown that systems for speech translation (ST) -- similarly to automatic speech recognition (ASR) -- poorly handle person names. This shortcoming does not only lead…

cs.CL20223 cited

Under the Morphosyntactic Lens: A Multifaceted Evaluation of Gender Bias in Speech Translation

Beatrice Savoldi, Marco Gaido, Luisa Bentivogli +2

Gender bias is largely recognized as a problematic phenomenon affecting language technologies, with recent studies underscoring that it might surface differently across languages.…

cs.CL2021

Is "moby dick" a Whale or a Bird? Named Entities and Terminology in Speech Translation

Marco Gaido, Susana Rodríguez, Matteo Negri +2

Automatic translation systems are known to struggle with rare words. Among these, named entities (NEs) and domain-specific terms are crucial, since errors in their translation can…