9 papers · 1 filter
SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
Mohamed Nabih Ali, Daniele Falavigna, Alessio Brutti
Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its application to large-scale speech…
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
Lorenzo Concina, Seraphina Fong, Marco Matassoni +1
Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for ta…
LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering
Nina Hosseini-Kivanani, Marco Matassoni, Alessio Brutti
Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM methods in low-resource settings.…
MLMA: Towards Multilingual ASR With Mamba-based Architectures
Mohamed Nabih Ali, Daniele Falavigna, Alessio Brutti
Multilingual automatic speech recognition (ASR) remains a challenging task, especially when balancing performance across high- and low-resource languages. Recent advances in sequen…
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
Maxence Lasbordes, Daniele Falavigna, Alessio Brutti
The ability to dynamically adjust the computational load of neural models during inference in a resource aware manner is crucial for on-device processing scenarios, characterised b…
FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
Sara Papi, Marco Gaido, Luisa Bentivogli +6
The development of speech foundation models (SFMs) like Whisper and SeamlessM4T has significantly advanced the field of speech processing. However, their closed nature--with inacce…