6 papers
SemanticAudio: Audio Generation and Editing in Semantic Space
Zheqi Dai, Guangyan Zhang, Haolin He +5
In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, exis…
One-Step Token-to-Waveform Generation with MeanFlow in Latent Space
Zheqi Dai, Guangyan Zhang, Zhen Ye +5
Neural audio codecs are central to modern LLM-based Text-to-Speech (TTS) and multimodal systems. As low-bitrate semantic codecs gain prominence, the Token-to-Waveform (Token2Wav) d…
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
Jingyu Li, Guangyan Zhang, Zhen Ye +1
Audio codecs are a critical component of modern speech generation systems. This paper introduces a low-bitrate, multi-scale residual codec that encodes speech into four distinct st…
CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
Aemon Yat Fei Chiu, Jingyu Li, Yusheng Tian +2
This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Ele…
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Jialong Zuo, Guangyan Zhang, Minghui Fang +5
Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms int…
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
Zhuangfei Cheng, Guangyan Zhang, Zehai Tu +6
Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, t…