papers

Publications (12)

cs.CL2019

Scalable Multi Corpora Neural Language Models for ASR

Anirudh Raju, Denis Filimonov, Gautam Tiwari +2

Neural language models (NLM) have been shown to outperform conventional n-gram language models by a substantial margin in Automatic Speech Recognition (ASR) and other tasks. There…

cs.CL2021

Domain-aware Neural Language Models for Speech Recognition

Linda Liu, Yile Gu, Aditya Gourav +5

As voice assistants become more ubiquitous, they are increasingly expected to support and perform well on a wide variety of use-cases across different domains. We present a domain-…

cs.CL2020

Neural Composition: Learning to Generate from Multiple Models

Denis Filimonov, Ravi Teja Gadde, Ariya Rastrow

Decomposing models into multiple components is critically important in many applications such as language modeling (LM) as it enables adapting individual components separately and…

cs.CL2024

Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition

Yu Yu, Chao-Han Huck Yang, Tuan Dinh +10

The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-c…

cs.CL2021

Personalization Strategies for End-to-End Speech Recognition Systems

Aditya Gourav, Linda Liu, Ankur Gandhe +9

The recognition of personalized content, such as contact names, remains a challenging problem for end-to-end speech recognition systems. In this work, we demonstrate how first and…

cs.CL2023

Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition

Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen +15

We propose a neural language modeling system based on low-rank adaptation (LoRA) for speech recognition output rescoring. Although pretrained language models (LMs) like BERT have s…

cs.CL2020

Improving accuracy of rare words for RNN-Transducer through unigram shallow fusion

Vijay Ravi, Yile Gu, Ankur Gandhe +5

End-to-end automatic speech recognition (ASR) systems, such as recurrent neural network transducer (RNN-T), have become popular, but rare word remains a challenge. In this paper, w…

cs.CL2021

Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Chao-Han Huck Yang, Linda Liu, Ankur Gandhe +4

End-to-end automatic speech recognition (ASR) systems are increasingly popular due to their relative architectural simplicity and competitive performance. However, even though the…

cs.AI2025

The Amazon Nova Family of Models: Technical Report and Model Card

Amazon AGI, Aaron Langford, Aayush Shah +783

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…

cs.CL2020

Neural Machine Translation For Paraphrase Generation

Alex Sokolov, Denis Filimonov

Training a spoken language understanding system, as the one in Alexa, typically requires a large human-annotated corpus of data. Manual annotations are expensive and time consuming…

eess.AS2023

PROCTER: PROnunciation-aware ConTextual adaptER for personalized speech recognition in neural transducers

Rahul Pandey, Roger Ren, Qi Luo +7

End-to-End (E2E) automatic speech recognition (ASR) systems used in voice assistants often have difficulties recognizing infrequent words personalized to the user, such as names an…

eess.AS2023

Streaming Speech-to-Confusion Network Speech Recognition

Denis Filimonov, Prabhat Pandey, Ariya Rastrow +2

In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-en…