5 papers
When De-noising Hurts: A Systematic Study of Speech Enhancement Effects on Modern Medical ASR Systems
Sujal Chondhekar, Vasanth Murukuri, Rushabh Vasani +8
Speech enhancement methods are commonly believed to improve the performance of automatic speech recognition (ASR) in noisy environments. However, the effectiveness of these techniq…
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
Shashi Kumar, Srikanth Madikeri, Esaú Villatoro-Tello +8
Token-based multitasking frameworks like TokenVerse require all training utterances to have labels for all tasks, hindering their ability to leverage partially annotated datasets a…
Unifying Global and Near-Context Biasing in a Single Trie Pass
Iuliia Thorbecke, Esaú Villatoro-Tello, Juan Zuluaga-Gomez +9
Despite the success of end-to-end automatic speech recognition (ASR) models, challenges persist in recognizing rare, out-of-vocabulary words - including named entities (NE) - and i…
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
Shashi Kumar, Srikanth Madikeri, Juan Zuluaga-Gomez +6
In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent pr…
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
Iuliia Thorbecke, Juan Zuluaga-Gomez, Esaú Villatoro-Tello +6
The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (T…