3 papers
eess.AS2025
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
Jingyu Li, Guangyan Zhang, Zhen Ye +1
Audio codecs are a critical component of modern speech generation systems. This paper introduces a low-bitrate, multi-scale residual codec that encodes speech into four distinct st…
cs.CL2025
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Jialong Zuo, Guangyan Zhang, Minghui Fang +5
Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms int…
eess.AS2025
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
Zhuangfei Cheng, Guangyan Zhang, Zehai Tu +6
Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, t…