collaborators

7 papers

cs.SD2026

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

Yue Heng Yeo, Haoyang Li, Yizhou Peng +6

Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs for training. Although synthetic data augme…

eess.AS2026

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

Haoyang Li, Xuyi Zhuang, Azmat Adnan +6

Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenT…

cs.CL2026

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

Shreyas Gopal, Donghang Wu, Ashutosh Anshul +5

Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficult to train with supervised fine…

eess.AS2026

Improving Code-Switching Speech Recognition with TTS Data Augmentation

Yue Heng Yeo, Yuchen Hu, Shreyas Gopal +3

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explo…

cs.CV2025

Next-Frame Feature Prediction for Multimodal Deepfake Detection and Temporal Localization

Ashutosh Anshul, Shreyas Gopal, Deepu Rajan +1

Recent multimodal deepfake detection methods designed for generalization conjecture that single-stage supervised training struggles to generalize across unseen manipulations and da…

cs.CL2025

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

Shreyas Gopal, Ashutosh Anshul, Haoyang Li +3

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for…