collaborators

5 papers

eess.AS2026

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

Zhisheng Zheng, Xiaohang Sun, Zhu Liu +5

Recent Text-To-Speech (TTS) systems have achieved strong naturalness and zero-shot voice cloning performance, but fine-grained control of expressive speech at the word or phoneme l…

eess.AS2025

VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing

Zhisheng Zheng, Puyuan Peng, Anuj Diwan +5

We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot Text-to-Speech (TTS) synthesis across 11 languages:…

eess.AS2025

RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speech

Zhisheng Zheng, Xiaohang Sun, Tuan Dinh +8

End-to-end speech-to-speech translation (S2ST) systems typically struggle with a critical data bottleneck: the scarcity of parallel speech-to-speech corpora. To overcome this, we i…

eess.AS2025

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

Puyuan Peng, Shang-Wen Li, Abdelrahman Mohamed +1

We present VoiceStar, the first zero-shot TTS model that achieves both output duration control and extrapolation. VoiceStar is an autoregressive encoder-decoder neural codec langua…

eess.AS2025

Scaling Rich Style-Prompted Text-to-Speech Datasets

Anuj Diwan, Zhisheng Zheng, David Harwath +1

We introduce Paralinguistic Speech Captions (ParaSpeechCaps), a large-scale dataset that annotates speech utterances with rich style captions. While rich abstract tags (e.g. guttur…