activity
20242026
most citedXLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models

1 citations · 1 across the 22 of their papers we have counts for

collaborators
Showing 2024Show all

5 papers · 1 filter

cs.CL2024

Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward

Shashi Kumar, Iuliia Thorbecke, Sergio Burdisso +7

Recent research has demonstrated that training a linear connector between speech foundation encoders and large language models (LLMs) enables this architecture to achieve strong AS…

cs.CL2024

Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper

Iuliia Thorbecke, Juan Zuluaga-Gomez, Esaú Villatoro-Tello +6

The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (T…

cs.CL2024

Unifying Global and Near-Context Biasing in a Single Trie Pass

Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9

Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…

cs.CL2024

TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR

Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6

In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…

eess.AS2024

XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models

Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +5

Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data. However, popular pretr…