66 citations · 68 across the 11 of their papers we have counts for
6 papers · 1 filter
Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, Rohit Prabhavalkar +5
Neural contextual biasing effectively improves automatic speech recognition (ASR) for crucial phrases within a speaker's context, particularly those that are infrequent in the trai…
Using Text Injection to Improve Recognition of Personal Identifiers in Speech
Yochai Blau, Rohan Agrawal, Lior Madmony +7
Accurate recognition of specific categories, such as persons' names, dates or other identifiers is critical in many Automatic Speech Recognition (ASR) applications. As these catego…
Understanding Shared Speech-Text Representations
Gary Wang, Kyle Kastner, Ankur Bapna +4
Recently, a number of approaches to train speech models by incorpo-rating text into end-to-end models have been developed, with Mae-stro advancing state-of-the-art automatic speech…
Robust Knowledge Distillation from RNN-T Models With Noisy Training Labels Using Full-Sum Loss
Mohammad Zeineldeen, Kartik Audhkhasi, Murali Karthick Baskar +1
This work studies knowledge distillation (KD) and addresses its constraints for recurrent neural network transducer (RNN-T) models. In hard distillation, a teacher model transcribe…
Invariant Representations for Noisy Speech Recognition
Dmitriy Serdyuk, Kartik Audhkhasi, Philémon Brakel +3
Modern automatic speech recognition (ASR) systems need to be robust under acoustic variability arising from environmental, speaker, channel, and recording conditions. Ensuring such…
Diverse Embedding Neural Network Language Models
Kartik Audhkhasi, Abhinav Sethy, Bhuvana Ramabhadran
We propose Diverse Embedding Neural Network (DENN), a novel architecture for language models (LMs). A DENNLM projects the input word history vector onto multiple diverse low-dimens…