4 papers
Improving Code-Switching Speech Recognition with TTS Data Augmentation
Yue Heng Yeo, Yuchen Hu, Shreyas Gopal +3
Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explo…
Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization
Ashutosh Anshul, Shreyas Gopal, Deepu Rajan +1
Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and da…
Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR
Shreyas Gopal, Ashutosh Anshul, Haoyang Li +3
Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for…
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
Alan Dao, Dinh Bach Vu, Huy Hoang Ha +6
The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of spee…