activity
20192024
most citedA Survey of Breast Cancer Screening Techniques: Thermography and Electrical Impedance Tomography

62 citations · 154 across the 27 of their papers we have counts for

collaborators

27 papers

cs.CL2024

Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper

Iuliia Thorbecke, Juan Zuluaga-Gomez, Esaú Villatoro-Tello +6

The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (T…

cs.CL2024

Unifying Global and Near-Context Biasing in a Single Trie Pass

Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9

Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…

cs.CL2024

TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR

Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6

In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…

eess.AS2024★ 1 cited

XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models

Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +5

Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data. However, popular pretr…

cs.LG2024★ 3 cited

Open-Source Conversational AI with SpeechBrain 1.0

Mirco Ravanelli, Titouan Parcollet, Adel Moumen +30

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker re…

cs.CL2023

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu +5

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversat…