6 papers
Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech
Omnilingual SONAR Team, João Maria Janeiro, Pere-LluÃs Huguet Cabot +17
Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstream quality for stronger alignment, limiting their adoption. We introduce OmniSO…
Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval
João Maria Janeiro, Mathurin Videau, Andrea Caciolai +3
Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically…
Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders
Belen Alastruey, João Maria Janeiro, Alexandre Allauzen +3
In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and…
Large Concept Models: Language Modeling in a Sentence Representation Space
LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18
LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…
FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data
Haoran Sun, Renren Jin, Shaoyang Xu +10
Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource lan…
MEXMA: Token-level objectives improve sentence representations
João Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari +1
Current pre-trained cross-lingual sentence encoders approaches use sentence-level objectives only. This can lead to loss of information, especially for tokens, which then degrades…