3 papers
eess.AS2025
SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model
Kaidi Wang, Yi He, Wenhao Guan +8
Video dubbing aims to generate high-fidelity speech that is precisely temporally aligned with the visual content. Existing methods still suffer from limitations in speech naturalne…
cs.SD2025
AVENet: Disentangling Features by Approximating Average Features for Voice Conversion
Wenyu Wang, Yiquan Zhou, Jihua Zhu +3
Voice conversion (VC) has made progress in feature disentanglement, but it is still difficult to balance timbre and content information. This paper evaluates the pre-trained model…
cs.SD2025
SYKI-SVC: Advancing Singing Voice Conversion with Post-Processing Innovations and an Open-Source Professional Testset
Yiquan Zhou, Wenyu Wang, Hongwu Ding +4
Singing voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this p…