collaborators

5 papers

cs.CL2023

Retrieve and Copy: Scaling ASR Personalization to Large Catalogs

Sai Muralidhar Jayanthi, Devang Kulshreshtha, Saket Dingliwal +2

Personalization of automatic speech recognition (ASR) models is a widely studied topic because of its many practical applications. Most recently, attention-based contextual biasing…

cs.SD2023

Generalized zero-shot audio-to-intent classification

Veera Raghavendra Elluru, Devang Kulshreshtha, Rohit Paturi +2

Spoken language understanding systems using audio-only data are gaining popularity, yet their ability to handle unseen intents remains limited. In this study, we propose a generali…

cs.SD2023

Masked Audio Text Encoders are Effective Multi-Modal Rescorers

Jinglun Cai, Monica Sunkara, Xilai Li +3

Masked Language Models (MLMs) have proven to be effective for second-pass rescoring in Automatic Speech Recognition (ASR) systems. In this work, we propose Masked Audio Text Encode…

eess.AS2023

Mask The Bias: Improving Domain-Adaptive Generalization of CTC-based ASR with Internal Language Model Estimation

Nilaksh Das, Monica Sunkara, Sravan Bodapati +4

End-to-end ASR models trained on large amount of data tend to be implicitly biased towards language semantics of the training data. Internal language model estimation (ILME) has be…

eess.AS2023

Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR

Xilai Li, Goeric Huybrechts, Srikanth Ronanki +2

Recently, there has been an increasing interest in unifying streaming and non-streaming speech recognition models to reduce development, training and deployment cost. The best-know…