activity
20182026
most citedStreaming Language Identification using Combination of Acoustic Representations and ASR Hypotheses

9 citations · 40 across the 31 of their papers we have counts for

collaborators
Showing 2024Show all

8 papers · 1 filter

eess.AS2024

CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing

Yen-Ju Lu, Jing Liu, Thomas Thebaud +4

We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to s…

eess.AS2024

Speech Recognition Rescoring with Large Speech-Text Foundation Models

Prashanth Gurunath Shivakumar, Jari Kolehmainen, Aditya Gourav +4

Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often…

eess.AS2024

An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems

Hitesh Tulsiani, David M. Chan, Shalini Ghosh +6

Dialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) system…

cs.CL2024

Multi-Modal Retrieval For Large Language Model Based Speech Recognition

Jari Kolehmainen, Aditya Gourav, Prashanth Gurunath Shivakumar +5

Retrieval is a widely adopted approach for improving language models leveraging external information. As the field moves towards multi-modal large language models, it is important…

cs.CL2024

Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition

Yu Yu, Chao-Han Huck Yang, Tuan Dinh +10

The use of low-rank adaptation (LoRA) with frozen pretrained language models (PLMs) has become increasing popular as a mainstream, resource-efficient modeling approach for memory-c…

eess.AS20243 cited

Two-pass Endpoint Detection for Speech Recognition

Anirudh Raju, Aparna Khare, Di He +9

Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accur…