collaborators

8 papers

cs.SD2025

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration

Yanru Huo, Ziyue Jiang, Zuoli Tang +2

While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech…

cs.CL2025

Entropy-based Coarse and Compressed Semantic Speech Representation Learning

Jialong Zuo, Guangyan Zhang, Minghui Fang +5

Discrete speech representation learning has recently attracted increasing interest in both acoustic and semantic modeling. Existing approaches typically encode 16 kHz waveforms int…

eess.AS2025

ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Shengpeng Ji, Qian Chen, Wen Wang +8

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style.…

eess.AS2025

Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

Shengpeng Ji, Minghui Fang, Jialong Zuo +5

In recent years, large language models have achieved significant success in generative tasks related to speech, audio, music, and other signal domains. A crucial element of these m…

eess.AS2025

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching

Jialong Zuo, Shengpeng Ji, Minghui Fang +7

Zero-Shot Voice Conversion (VC) aims to transform the source speaker's timbre into an arbitrary unseen one while retaining speech content. Most prior work focuses on preserving the…

cs.SD2025

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Jialong Zuo, Shengpeng Ji, Minghui Fang +8

This paper introduces PFlow-VC, a conditional flow matching voice conversion model that leverages fine-grained discrete pitch tokens and target speaker prompt information for expre…