5 papers · 1 filter
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
Lorenzo Concina, Seraphina Fong, Marco Matassoni +1
Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for ta…
LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering
Nina Hosseini-Kivanani, Marco Matassoni, Alessio Brutti
Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM methods in low-resource settings.…
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
Sara Papi, Marco Gaido, Luisa Bentivogli +6
The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inacce…
The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence
Marco Gaido, Sara Papi, Luisa Bentivogli +6
Training large-scale models presents challenges not only in terms of resource requirements but also in terms of their convergence. For this reason, the learning rate (LR) is often…
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
Marco Gaido, Sara Papi, Luisa Bentivogli +6
The rise of foundation models (FMs), coupled with regulatory efforts addressing their risks and impacts, has sparked significant interest in open-source models. However, existing s…