2 papers
cs.CL2026
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
Franck Signe, Hippolyte Pilchen, François Yvon +1
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, t…
cs.CL2022
Joint Generation of Captions and Subtitles with Dual Decoding
Jitao Xu, François Buet, Josep Crego +2
As the amount of audio-visual content increases, the need to develop automatic captioning and subtitling solutions to match the expectations of a growing international audience app…