activity
20242026
collaborators

5 papers

eess.AS2026

BaldWhisper: Faster Whisper with Head Shearing and Layer Merging

Yaya Sy, Christophe Cerisara, Irina Illina

Pruning large pre-trained transformers in a data-scarce scenario is challenging, as it often requires massive retraining data to recover performance. For instance, Distill-Whisper…

cs.CL2026

Cross-lingual Matryoshka Representation Learning across Speech and Text

Yaya Sy, Dioula Doucouré, Christophe Cerisara +1

Speakers of under-represented languages face both a language barrier, as most online knowledge is in a few dominant languages, and a modality barrier, since information is largely…

cs.CL2025

Speech Language Models for Under-Represented Languages: Insights from Wolof

Yaya Sy, Dioula Doucouré, Christophe Cerisara +1

We present our journey in training a speech language model for Wolof, an underrepresented language spoken in West Africa, and share key insights. We first emphasize the importance…

cs.CL2025

The Lucie-7B LLM and the Lucie Training Dataset: Open resources for multilingual language generation

Olivier Gouvert, Julie Hunter, Jérôme Louradour +6

We present both the Lucie Training Dataset and the Lucie-7B foundation model. The Lucie Training Dataset is a multilingual collection of textual corpora centered around French and…

cs.LG2024

Lillama: Large Language Models Compression via Low-Rank Feature Distillation

Yaya Sy, Christophe Cerisara, Irina Illina

Current LLM structured pruning methods typically involve two steps: (1) compression with calibration data and (2) costly continued pretraining on billions of tokens to recover lost…