3 papers
eess.AS2026
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
Mengqi Wang, Zhan Liu, Zengrui Jin +3
Diffusion-based large language models (DLLMs) have recently attracted growing interest as an alternative to autoregressive decoders. In this work, we present an empirical study on…
cs.SD2025
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
Qiuming Zhao, Guangzhi Sun, Chao Zhang
Language diversity presents a significant challenge in speech-to-text (S2T) tasks, such as automatic speech recognition and translation. Traditional multi-lingual multi-task traini…
cs.SD2025
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
Yiyang Zhao, Shuai Wang, Guangzhi Sun +4
Short-utterance speaker verification presents significant challenges due to the limited information in brief speech segments, which can undermine accuracy and reliability. Recently…