2 citations · 4 across the 3 of their papers we have counts for
9 papers
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
Pooneh Mousavi, Jarod Duret, Salah Zaiem +4
Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paraling…
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
Luca Zampierin, Ghouthi Boukli Hacene, Bac Nguyen +1
Self-supervised learning (SSL) has achieved remarkable success across various speech-processing tasks. To enhance its efficiency, previous works often leverage the use of compressi…
Focal Modulation Networks for Interpretable Sound Classification
Luca Della Libera, Cem Subakan, Mirco Ravanelli
The increasing success of deep neural networks has raised concerns about their inherent black-box nature, posing challenges related to interpretability and trust. While there has b…
Bayesian Deep Learning for Remaining Useful Life Estimation via Stein Variational Gradient Descent
Luca Della Libera, Jacopo Andreoli, Davide Dalle Pezze +2
A crucial task in predictive maintenance is estimating the remaining useful life of physical systems. In the last decade, deep learning has improved considerably upon traditional m…
Are LLMs Robust for Spoken Dialogues?
Seyed Mahed Mousavi, Gabriel Roccabruna, Simone Alghisi +3
Large Pre-Trained Language Models have demonstrated state-of-the-art performance in different downstream tasks, including dialogue state tracking and end-to-end response generation…
CL-MASR: A Continual Learning Benchmark for Multilingual ASR
Luca Della Libera, Pooneh Mousavi, Salah Zaiem +2
Modern multilingual automatic speech recognition (ASR) systems like Whisper have made it possible to transcribe audio in multiple languages with a single model. However, current st…