4 citations · 7 across the 5 of their papers we have counts for
5 papers
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
Chang Zeng, Chunhui Wang, Xiaoxiao Miao +3
It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural voco…
Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
Chang Zeng, Xiaoxiao Miao, Xin Wang +2
In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel misma…
Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network
Chunhui Wang, Chang Zeng, Xing He
XiaoiceSing is a singing voice synthesis (SVS) system that aims at generating 48kHz singing voices. However, the mel-spectrogram generated by it is over-smoothing in middle- and hi…
Deep Spectro-temporal Artifacts for Detecting Synthesized Speech
Xiaohui Liu, Meng Liu, Lin Zhang +7
The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of trac…
Exploring Deep Learning for Joint Audio-Visual Lip Biometrics
Meng Liu, Longbiao Wang, Kong Aik Lee +3
Audio-visual (AV) lip biometrics is a promising authentication technique that leverages the benefits of both the audio and visual modalities in speech communication. Previous works…