13 citations · 16 across the 5 of their papers we have counts for
5 papers
Retrieval-Augmented Audio Deepfake Detection
Zuheng Kang, Yayun He, Botao Zhao +4
With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growi…
VoiceExtender: Short-utterance Text-independent Speaker Verification with Guided Diffusion Model
Yayun He, Zuheng Kang, Jianzong Wang +2
Speaker verification (SV) performance deteriorates as utterances become shorter. To this end, we propose a new architecture called VoiceExtender which provides a promising solution…
SVVAD: Personal Voice Activity Detection for Speaker Verification
Zuheng Kang, Jianzong Wang, Junqing Peng +1
Voice activity detection (VAD) improves the performance of speaker verification (SV) by preserving speech segments and attenuating the effects of non-speech. However, this scheme i…
Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification
Zuheng Kang, Yayun He, Jianzong Wang +3
Data-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promisin…
SpeechEQ: Speech Emotion Recognition based on Multi-scale Unified Datasets and Multitask Learning
Zuheng Kang, Junqing Peng, Jianzong Wang +1
Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a…