papers

Publications (21)

cs.CL2024

Large Concept Models: Language Modeling in a Sentence Representation Space

LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18

LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…

cs.CL2025

Improving Language and Modality Transfer in Translation by Character-level Modeling

Ioannis Tsiamas, David Dale, Marta R. Costa-jussÃ

Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requi…

cs.CL2025

BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

The Omnilingual MT Team, Pierre Andrews, Mikel Artetxe +14

BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages…

cs.CL2023

Exploring Methods for Cross-lingual Text Style Transfer: The Case of Text Detoxification

Daryna Dementieva, Daniil Moskovskiy, David Dale +1

Text detoxification is the task of transferring the style of text from toxic to neutral. While here are approaches yielding promising results in monolingual setup, e.g., (Dale et a…

cs.CL2021

Text Detoxification using Large Pre-trained Neural Models

David Dale, Anton Voronov, Daryna Dementieva +4

We present two novel unsupervised methods for eliminating toxicity in text. Our first method combines two recent ideas: (1) guidance of the generation process with small style-cond…

cs.CL2024

SpeechAlign: a Framework for Speech Translation Alignment Evaluation

Belen Alastruey, Aleix Sant, Gerard I. Gállego +2

Speech-to-Speech and Speech-to-Text translation are currently dynamic areas of research. In our commitment to advance these fields, we present SpeechAlign, a framework designed to…