activity
20212024
most citedDomain-aware Neural Language Models for Speech Recognition

2 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2024

Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Guan-Ting Lin, Prashanth Gurunath Shivakumar, Aditya Gourav +4

While textless Spoken Language Models (SLMs) have shown potential in end-to-end speech-to-speech modeling, they still lag behind text-based Large Language Models (LLMs) in terms of…

cs.CL2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

Jari Kolehmainen, Aditya Gourav, Prashanth Gurunath Shivakumar +5

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important…

cs.CL2023

Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition

Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen +15

We propose a neural language modeling system based on low-rank adaptation (LoRA) for speech recognition output rescoring. Although pretrained language models (LMs) like BERT have s…

cs.CL2023

On-the-fly Text Retrieval for End-to-End ASR Adaptation

Bolaji Yusuf, Aditya Gourav, Ankur Gandhe +1

End-to-end speech recognition models are improved by incorporating external text sources, typically by fusion with an external language model. Such language models have to be retra…

cs.CL20212 cited

Domain-aware Neural Language Models for Speech Recognition

Linda Liu, Yile Gu, Aditya Gourav +5

As voice assistants become more ubiquitous, they are increasingly expected to support and perform well on a wide variety of use-cases across different domains. We present a domain-…

cs.CL2021

Personalization Strategies for End-to-End Speech Recognition Systems

Aditya Gourav, Linda Liu, Ankur Gandhe +9

The recognition of personalized content, such as contact names, remains a challenging problem for end-to-end speech recognition systems. In this work, we demonstrate how first and…