Publications (21)
Large Concept Models: Language Modeling in a Sentence Representation Space
LCM team, Loïc Barrault, Paul-Ambroise Duquenne +18
LLMs have revolutionized the field of artificial intelligence and have emerged as the de-facto tool for many tasks. The current established technology of LLMs is to process input a…
Improving Language and Modality Transfer in Translation by Character-level Modeling
Ioannis Tsiamas, David Dale, Marta R. Costa-jussÃ
Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requi…
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
The Omnilingual MT Team, Pierre Andrews, Mikel Artetxe +14
BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages…
Exploring Methods for Cross-lingual Text Style Transfer: The Case of Text Detoxification
Daryna Dementieva, Daniil Moskovskiy, David Dale +1
Text detoxification is the task of transferring the style of text from toxic to neutral. While here are approaches yielding promising results in monolingual setup, e.g., (Dale et a…
Text Detoxification using Large Pre-trained Neural Models
David Dale, Anton Voronov, Daryna Dementieva +4
We present two novel unsupervised methods for eliminating toxicity in text. Our first method combines two recent ideas: (1) guidance of the generation process with small style-cond…
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
Belen Alastruey, Aleix Sant, Gerard I. Gállego +2
Speech-to-Speech and Speech-to-Text translation are currently dynamic areas of research. In our commitment to advance these fields, we present SpeechAlign, a framework designed to…