2 papers
cs.CL2024
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
Yash Jain, David Chan, Pranav Dheram +4
Recent advances in machine learning have demonstrated that multi-modal pre-training can improve automatic speech recognition (ASR) performance compared to randomly initialized mode…
cs.CL2024
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
Jinhan Wang, Long Chen, Aparna Khare +6
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM).…