activity
20242026
collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2026

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Peijie Chen, Wenhao Guan, Weijie Wu +7

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audi…

cs.SD2025

XMUspeech Systems for the ASVspoof 5 Challenge

Wangjie Li, Xingjia Xie, Yishuang Li +5

In this paper, we present our submitted XMUspeech systems to the speech deepfake detection track of the ASVspoof 5 Challenge. Compared to previous challenges, the audio duration in…

cs.SD2025

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization

Pengyu Ren, Wenhao Guan, Kaidi Wang +3

In recent years, diffusion-based generative models have demonstrated remarkable performance in speech conversion, including Denoising Diffusion Probabilistic Models (DDPM) and othe…

cs.SD2025

A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement

Shenghui Lu, Hukai Huang, Jinanglong Yao +3

This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins…

cs.SD2025

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

Peijie Chen, Wenhao Guan, Kaidi Wang +4

Neural speech codecs are essential for advancing text-to-speech (TTS) systems. With the recent success of large language models in text generation, developing high-quality speech t…

cs.SD2025

Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion

Kaidi Wang, Wenhao Guan, Ziyue Jiang +5

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speak…