7 citations · 10 across the 4 of their papers we have counts for
4 papers
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
Jinhan Wang, Long Chen, Aparna Khare +6
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM).…
Two-pass Endpoint Detection for Speech Recognition
Anirudh Raju, Aparna Khare, Di He +9
Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accur…
ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale
Gopinath Chennupati, Milind Rao, Gurpreet Chadha +11
Incremental learning is one paradigm to enable model building and updating at scale with streaming data. For end-to-end automatic speech recognition (ASR) tasks, the absence of hum…
Attentive Contextual Carryover for Multi-Turn End-to-End Spoken Language Understanding
Kai Wei, Thanh Tran, Feng-Ju Chang +8
Recent years have seen significant advances in end-to-end (E2E) spoken language understanding (SLU) systems, which directly predict intents and slots from spoken audio. While dialo…