activity
20232026
most citedLarge Concept Models: Language Modeling in a Sentence Representation Space

10 citations · 12 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

João Maria Janeiro, Mathurin Videau, Andrea Caciolai +3

Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically…

cs.CL2026

Omnilingual MT: Machine Translation for 1,600 Languages

Omnilingual MT Team, Belen Alastruey, Niyati Bafna +29

High-quality machine translation (MT) can scale to hundreds of languages, setting a high bar for multilingual systems. However, compared to the world's 7,000 languages, current sys…

cs.CL2025

Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders

Belen Alastruey, João Maria Janeiro, Alexandre Allauzen +3

In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and…

cs.CL202410 cited

Large Concept Models: Language Modeling in a Sentence Representation Space

LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18

LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…

cs.CL2024

MEXMA: Token-level objectives improve sentence representations

João Maria Janeiro, Benjamin Piwowarski, Patrick Gallinari +1

Current pre-trained cross-lingual sentence encoders approaches use sentence-level objectives only. This can lead to loss of information, especially for tokens, which then degrades…