62 citations · 154 across the 27 of their papers we have counts for
27 papers
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
Iuliia Thorbecke, Juan Zuluaga-Gomez, Esaú Villatoro-Tello +6
The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (T…
Unifying Global and Near-Context Biasing in a Single Trie Pass
Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9
Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6
In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +5
Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data. However, popular pretr…
Open-Source Conversational AI with SpeechBrain 1.0
Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…
End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation
Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu +5
Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversat…