activity
20212024
most citedExploring Deep Learning for Joint Audio-Visual Lip Biometrics

4 citations · 7 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself

Chang Zeng, Chunhui Wang, Xiaoxiao Miao +3

It is challenging to accelerate the training process while ensuring both high-quality generated voices and acceptable inference speed. In this paper, we propose a novel neural voco…

eess.AS2024

Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches

Chang Zeng, Xiaoxiao Miao, Xin Wang +2

In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel misma…

eess.AS20221 cited

Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network

Chunhui Wang, Chang Zeng, Xing He

XiaoiceSing is a singing voice synthesis (SVS) system that aims at generating 48kHz singing voices. However, the mel-spectrogram generated by it is over-smoothing in middle- and hi…

cs.SD20222 cited

Deep Spectro-temporal Artifacts for Detecting Synthesized Speech

Xiaohui Liu, Meng Liu, Lin Zhang +7

The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of trac…

cs.MM20214 cited

Exploring Deep Learning for Joint Audio-Visual Lip Biometrics

Meng Liu, Longbiao Wang, Kong Aik Lee +3

Audio-visual (AV) lip biometrics is a promising authentication technique that leverages the benefits of both the audio and visual modalities in speech communication. Previous works…