most citedMETTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

1 citations · 3 across the 14 of their papers we have counts for

collaborators

14 papers

cs.SD2024

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls

Zihao Chen, Haomin Zhang, Xinhan Di +10

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-…

cs.SD2024

CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion

Yuke Li, Xinfa Zhu, Hanzhao Li +6

Zero-shot voice conversion (VC) aims to convert the original speaker's timbre to any target speaker while keeping the linguistic content. Current mainstream zero-shot voice convers…

cs.SD2024

The NPU-HWC System for the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge

Dake Guo, Jixun Yao, Xinfa Zhu +6

This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a spee…

cs.SD2024

Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy

Linhan Ma, Xinfa Zhu, Yuanjun Lv +5

Zero-shot voice conversion (VC) aims to transform source speech into arbitrary unseen target voice while keeping the linguistic content unchanged. Recent VC methods have made signi…

eess.AS2024

Text-aware and Context-aware Expressive Audiobook Speech Synthesis

Dake Guo, Xinfa Zhu, Liumeng Xue +3

Recent advances in text-to-speech have significantly improved the expressiveness of synthetic speech. However, a major challenge remains in generating speech that captures the dive…

eess.AS2024

Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Hanzhao Li, Liumeng Xue, Haohan Guo +6

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid t…