most citedSoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation

2 citations · 4 across the 3 of their papers we have counts for

collaborators

9 papers

cs.SD2024

How Should We Extract Discrete Audio Tokens from Self-Supervised Models?

Pooneh Mousavi, Jarod Duret, Salah Zaiem +4

Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paraling…

eess.AS2024

SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning

Luca Zampierin, Ghouthi Boukli Hacene, Bac Nguyen +1

Self-supervised learning (SSL) has achieved remarkable success across various speech-processing tasks. To enhance its efficiency, previous works often leverage the use of compressi…

cs.SD2024

Focal Modulation Networks for Interpretable Sound Classification

Luca Della Libera, Cem Subakan, Mirco Ravanelli

The increasing success of deep neural networks has raised concerns about their inherent black-box nature, posing challenges related to interpretability and trust. While there has b…

cs.LG20243 cited

Bayesian Deep Learning for Remaining Useful Life Estimation via Stein Variational Gradient Descent

Luca Della Libera, Jacopo Andreoli, Davide Dalle Pezze +2

A crucial task in predictive maintenance is estimating the remaining useful life of physical systems. In the last decade, deep learning has improved considerably upon traditional m…

cs.CL2024

Are LLMs Robust for Spoken Dialogues?

Seyed Mahed Mousavi, Gabriel Roccabruna, Simone Alghisi +3

Large Pre-Trained Language Models have demonstrated state-of-the-art performance in different downstream tasks, including dialogue state tracking and end-to-end response generation…

cs.CL2023

CL-MASR: A Continual Learning Benchmark for Multilingual ASR

Luca Della Libera, Pooneh Mousavi, Salah Zaiem +2

Modern multilingual automatic speech recognition (ASR) systems like Whisper have made it possible to transcribe audio in multiple languages with a single model. However, current st…