3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
Yash Jain, David Chan, Pranav Dheram +4
Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized mode…
cs.CL2024
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
Jinhan Wang, Long Chen, Aparna Khare +6
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM).…
eess.AS2024★ 3 cited
Two-pass Endpoint Detection for Speech Recognition
Anirudh Raju, Aparna Khare, Di He +9
Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accur…