3 papers
eess.AS2026
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
Qibing Bai, Shuhao Shi, Shuai Wang +3
Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. I…
cs.SD2023
MC-SpEx: Towards Effective Speaker Extraction with Multi-Scale Interfusion and Conditional Speaker Modulation
Jun Chen, Wei Rao, Zilin Wang +5
The previous SpEx+ has yielded outstanding performance in speaker extraction and attracted much attention. However, it still encounters inadequate utilization of multi-scale inform…
cs.SD2022
Speech Enhancement with Intelligent Neural Homomorphic Synthesis
Shulin He, Wei Rao, Jinjiang Liu +5
Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a…