2 papers
cs.SD2024
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
Pooneh Mousavi, Jarod Duret, Salah Zaiem +4
Discrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing. Ideal audio tokens must preserve content, paraling…
cs.CL2023
CL-MASR: A Continual Learning Benchmark for Multilingual ASR
Luca Della Libera, Pooneh Mousavi, Salah Zaiem +2
Modern multilingual automatic speech recognition (ASR) systems like Whisper have made it possible to transcribe audio in multiple languages with a single model. However, current st…