activity
20242026
collaborators

12 papers

eess.AS2026

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

Ye-Xin Lu, Xin Wang, Yang Ai +3

Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…

eess.AS2026

DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis

Ye-Xin Lu, Yu Gu, Kun Wei +3

This paper presents DAIEN-TTS, a zero-shot text-to-speech (TTS) framework that enables ENvironment-aware synthesis through Disentangled Audio Infilling. By leveraging separate spea…

cs.SD2025

Universal Discrete-Domain Speech Enhancement

Fei Liu, Yang Ai, Ye-Xin Lu +3

In real-world scenarios, speech signals are inevitably corrupted by various types of interference, making speech enhancement (SE) a critical task for robust speech processing. Howe…

eess.AS2025

Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?

Hui-Peng Du, Yang Ai, Rui-Chen Zheng +2

Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, th…

eess.AS2025

Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising

Ye-Xin Lu, Hui-Peng Du, Fei Liu +2

Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech…

cs.SD2024

Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion

Yu-Fei Shi, Yang Ai, Ye-Xin Lu +2

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all…