activity
20232025
most citedSeamless: Multilingual Expressive and Streaming Speech Translation

41 citations · 42 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2025

BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

The Omnilingual MT Team, Pierre Andrews, Mikel Artetxe +14

BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages…

cs.CL2024

LCFO: Long Context and Long Form Output Dataset and Benchmarking

Marta R. Costa-jussà, Pierre Andrews, Mariano Coria Meglioli +10

This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across di…

cs.CL2024

On the Role of Speech Data in Reducing Toxicity Detection Bias

Samuel J. Bell, Mariano Coria Meglioli, Megan Richards +6

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity de…

cs.SD2024★ 1 cited

MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

Marta R. Costa-jussà, Mariano Coria Meglioli, Pierre Andrews +6

Research in toxicity detection in natural language processing for the speech modality (audio-based) is quite limited, particularly for languages other than English. To address thes…

cs.CL2023★ 41 cited

Seamless: Multilingual Expressive and Streaming Speech Translation

Seamless Communication, Loïc Barrault, Yu-An Chung +62

Large-scale automatic speech translation systems today lack key features that help machine-mediated communication feel seamless when compared to human-to-human dialogue. In this wo…