8 papers
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
Omnilingual SONAR Team, João Maria Janeiro, Pere-LluÃs Huguet Cabot +17
Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSO…
Omnilingual MT: Machine Translation for 1,600 Languages
Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29
High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current sys…
Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders
Belen Alastruey, João Maria Janeiro, Alexandre Allauzen +3
In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and…
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
Marta R. Costa-jussÃ, Bokai Yu, Pierre Andrews +7
We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the inters…
Large Concept Models: Language Modeling in a Sentence Representation Space
LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18
LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…
Unveiling the Role of Pretraining in Direct Speech Translation
Belen Alastruey, Gerard I. Gállego, Marta R. Costa-jussÃ
Direct speech-to-text translation systems encounter an important drawback in data scarcity. A common solution consists on pretraining the encoder on automatic speech recognition, h…