12 papers
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Ye-Xin Lu, Xin Wang, Yang Ai +3
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
Ye-Xin Lu, Yu Gu, Kun Wei +3
This paper presents DAIEN-TTS, a zero-shot text-to-speech (TTS) framework that enables ENvironment-aware synthesis through Disentangled Audio Infilling. By leveraging separate spea…
Universal Discrete-Domain Speech Enhancement
Fei Liu, Yang Ai, Ye-Xin Lu +3
In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. Howe…
Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
Hui-Peng Du, Yang Ai, Rui-Chen Zheng +2
Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, th…
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
Ye-Xin Lu, Hui-Peng Du, Fei Liu +2
Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech…
Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion
Yu-Fei Shi, Yang Ai, Ye-Xin Lu +2
We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all…