3 citations · 4 across the 4 of their papers we have counts for
4 papers
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
Yash Jain, David Chan, Pranav Dheram +4
Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized mode…
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
Jinhan Wang, Long Chen, Aparna Khare +6
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM).…
Two-pass Endpoint Detection for Speech Recognition
Anirudh Raju, Aparna Khare, Di He +9
Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accur…
Improving fairness for spoken language understanding in atypical speech with Text-to-Speech
Helin Wang, Venkatesh Ravichandran, Milind Rao +8
Spoken language understanding (SLU) systems often exhibit suboptimal performance in processing atypical speech, typically caused by neurological conditions and motor impairments. R…