most citedParalinguistics-Enhanced Large Language Modeling of Spoken Dialogue

1 citations · 1 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

Jari Kolehmainen, Aditya Gourav, Prashanth Gurunath Shivakumar +5

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important…

cs.CL2024

Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition

Yu Yu, Chao-Han Huck Yang, Tuan Dinh +10

The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-c…

cs.CL20241 cited

Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue

Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe +6

Large Language Models (LLMs) have demonstrated superior abilities in tasks such as chatting, reasoning, and question-answering. However, standard LLMs may ignore crucial paralingui…

cs.CL2024

Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks

Kevin Everson, Yile Gu, Huck Yang +10

In the realm of spoken language understanding (SLU), numerous natural language understanding (NLU) methodologies have been adapted by supplying large language models (LLMs) with tr…

eess.AS2023

Discriminative Speech Recognition Rescoring with Pre-trained Language Models

Prashanth Gurunath Shivakumar, Jari Kolehmainen, Yile Gu +3

Second pass rescoring is a critical component of competitive automatic speech recognition (ASR) systems. Large language models have demonstrated their ability in using pre-trained…

eess.AS2023

Personalization for BERT-based Discriminative Speech Recognition Rescoring

Jari Kolehmainen, Yile Gu, Aditya Gourav +4

Recognition of personalized content remains a challenge in end-to-end speech recognition. We explore three novel approaches that use personalized content in a neural rescoring step…