Publications (12)
Scalable Multi Corpora Neural Language Models for ASR
Anirudh Raju, Denis Filimonov, Gautam Tiwari +2
Neural language models (NLM) have been shown to outperform conventional n-gram language models by a substantial margin in Automatic Speech Recognition (ASR) and other tasks. There…
Domain-aware Neural Language Models for Speech Recognition
Linda Liu, Yile Gu, Aditya Gourav +5
As voice assistants become more ubiquitous, they are increasingly expected to support and perform well on a wide variety of use-cases across different domains. We present a domain-…
Neural Composition: Learning to Generate from Multiple Models
Denis Filimonov, Ravi Teja Gadde, Ariya Rastrow
Decomposing models into multiple components is critically important in many applications such as language modeling (LM) as it enables adapting individual components separately and…
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition
Yu Yu, Chao-Han Huck Yang, Tuan Dinh +10
The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-c…
Personalization Strategies for End-to-End Speech Recognition Systems
Aditya Gourav, Linda Liu, Ankur Gandhe +9
The recognition of personalized content, such as contact names, remains a challenging problem for end-to-end speech recognition systems. In this work, we demonstrate how first and…
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen +15
We propose a neural language modeling system based on low-rank adaptation (LoRA) for speech recognition output rescoring. Although pretrained language models (LMs) like BERT have s…
Improving accuracy of rare words for RNN-Transducer through unigram shallow fusion
Vijay Ravi, Yile Gu, Ankur Gandhe +5
End-to-end automatic speech recognition (ASR) systems, such as recurrent neural network transducer (RNN-T), have become popular, but rare word remains a challenge. In this paper, w…
Multi-task Language Modeling for Improving Speech Recognition of Rare Words
Chao-Han Huck Yang, Linda Liu, Ankur Gandhe +4
End-to-end automatic speech recognition (ASR) systems are increasingly popular due to their relative architectural simplicity and competitive performance. However, even though the…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
Neural Machine Translation For Paraphrase Generation
Alex Sokolov, Denis Filimonov
Training a spoken language understanding system, as the one in Alexa, typically requires a large human-annotated corpus of data. Manual annotations are expensive and time consuming…
PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers
Rahul Pandey, Roger Ren, Qi Luo +7
End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names an…
Streaming Speech-to-Confusion Network Speech Recognition
Denis Filimonov, Prabhat Pandey, Ariya Rastrow +2
In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-en…