19 citations · 19 across the 4 of their papers we have counts for
4 papers
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
Zuheng Kang, Yayun He, Jianzong Wang +2
Single-model systems often suffer from deficiencies in tasks such as speaker verification (SV) and image classification, relying heavily on partial prior knowledge during decision-…
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
Haobin Tang, Xulong Zhang, Ning Cheng +2
Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech…
Symbolic & Acoustic: Multi-domain Music Emotion Modeling for Instrumental Music
Kexin Zhu, Xulong Zhang, Jianzong Wang +2
Music Emotion Recognition involves the automatic identification of emotional elements within music tracks, and it has garnered significant attention due to its broad applicability…
TGAVC: Improving Autoencoder Voice Conversion with Text-Guided and Adversarial Training
Huaizhen Tang, Xulong Zhang, Jianzong Wang +4
Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excelle…