activity
20202026
most citedTowards Multi-Scale Style Control for Expressive Speech Synthesis

4 citations · 7 across the 8 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

Xiang Li, Yixuan Zhou, Jingran Xie +2

Neural speech codecs based on Vector-Quantized VAEs (VQ-VAEs) are core audio tokenizers for speech LLMs, yet their reconstruction fidelity is bottlenecked by quantization error. Mo…

cs.SD2023

A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis

Xianhao Wei, Jia Jia, Xiang Li +2

This paper explores predicting suitable prosodic features for fine-grained emotion analysis from the discourse-level text. To obtain fine-grained emotional prosodic features as pre…

cs.SD2023

CALM: Contrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis

Yi Meng, Xiang Li, Zhiyong Wu +6

To further improve the speaking styles of synthesized speeches, current text-to-speech (TTS) synthesis systems commonly employ reference speeches to stylize their outputs instead o…

cs.SD2023

Diverse and Expressive Speech Prosody Prediction with Denoising Diffusion Probabilistic Model

Xiang Li, Songxiang Liu, Max W. Y. Lam +3

Expressive human speech generally abounds with rich and flexible speech prosody variations. The speech prosody predictors in existing expressive speech synthesis methods mostly pro…

cs.SD20214 cited

Towards Multi-Scale Style Control for Expressive Speech Synthesis

Xiang Li, Changhe Song, Jingbei Li +3

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract…