3 papers
cs.SD2026
DASB - Discrete Audio and Speech Benchmark
Pooneh Mousavi, Jarod Duret, Darius Petermann +5
Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multimodal language models that can both gen…
cs.CL2025
In-domain SSL pre-training and streaming ASR
Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière +6
In this study, we investigate the benefits of domain-specific self-supervised pre-training for both offline and streaming ASR in Air Traffic Control (ATC) environments. We train BE…
cs.LG2024
Open-Source Conversational AI with SpeechBrain 1.0
Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…