collaborators

14 papers

cs.SD2024

Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation

Hongming Guo, Ruibo Fu, Yizhong Geng +9

Text-to-audio (TTA) model is capable of generating diverse audio from textual prompts. However, most mainstream TTA models, which predominantly rely on Mel-spectrograms, still face…

eess.AS2024

The FruitShell French synthesis system at the Blizzard 2023 Challenge

Xin Qi, Xiaopeng Wang, Zhiyong Wang +3

This paper presents a French text-to-speech synthesis system for the Blizzard Challenge 2023. The challenge consists of two tasks: generating high-quality speech from female speake…

cs.SD2024

Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0

Zhiyong Wang, Ruibo Fu, Zhengqi Wen +10

Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and in…

cs.SD2024

DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech

Xin Qi, Ruibo Fu, Zhengqi Wen +12

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have…

eess.AS2024

Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation

Chenxu Xiong, Ruibo Fu, Shuchen Shi +9

Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To addre…

cs.SD2024

EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech

Xin Qi, Ruibo Fu, Zhengqi Wen +10

In the current era of Artificial Intelligence Generated Content (AIGC), a Low-Rank Adaptation (LoRA) method has emerged. It uses a plugin-based approach to learn new knowledge with…