activity
20242026
collaborators

8 papers

cs.CL2026

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

Omnilingual SONAR Team, João Maria Janeiro, Pere-Lluís Huguet Cabot +17

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSO…

cs.CL2026

Omnilingual MT: Machine Translation for 1,600 Languages

Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29

High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current sys…

cs.CL2025

Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders

Belen Alastruey, João Maria Janeiro, Alexandre Allauzen +3

In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and…

cs.CL2024

2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset

Marta R. Costa-jussÃ, Bokai Yu, Pierre Andrews +7

We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 74 spoken languages at the inters…

cs.CL2024

Large Concept Models: Language Modeling in a Sentence Representation Space

LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18

LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…

cs.CL2024

Unveiling the Role of Pretraining in Direct Speech Translation

Belen Alastruey, Gerard I. Gállego, Marta R. Costa-jussÃ

Direct speech-to-text translation systems encounter an important drawback in data scarcity. A common solution consists on pretraining the encoder on automatic speech recognition, h…