6 papers
Unifying Global and Near-Context Biasing in a Single Trie Pass
Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9
Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…
Improving endpoint detection in end-to-end streaming ASR for conversational speech
Anandh C, Karthik Pandia Durai, Jeena Prakash +6
ASR endpointing (EP) plays a major role in delivering a good user experience in products supporting human or artificial agents in human-human/machine conversations. Transducer-base…
Performance evaluation of SLAM-ASR: The Good, the Bad, the Ugly, and the Way Forward
Shashi Kumar, Iuliia Thorbecke, Sergio Burdisso +7
Recent research has demonstrated that training a linear connector between speech foundation encoders and large language models (LLMs) enables this architecture to achieve strong AS…
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6
In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +5
Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data. However, popular pretr…
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
Iuliia Thorbecke, Juan Zuluaga-Gomez, Esaú Villatoro-Tello +6
The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (T…