5 papers
Omnilingual MT: Machine Translation for 1,600 Languages
Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29
High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current sys…
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
Omnilingual ASR team, Gil Keren, Artyom Kozhevnikov +30
Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages be…
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
The Omnilingual MT Team, Pierre Andrews, Mikel Artetxe +14
BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages…
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
Zheng-Xin Yong, Vineel Pratap, Michael Auli +1
To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We syste…
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
Marta R. Costa-jussÃ, Bokai Yu, Pierre Andrews +7
We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the inters…