activity
20222024
most citedRetrieval-Augmented Audio Deepfake Detection

13 citations · 16 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD202413 cited

Retrieval-Augmented Audio Deepfake Detection

Zuheng Kang, Yayun He, Botao Zhao +4

With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growi…

cs.SD2023

VoiceExtender: Short-utterance Text-independent Speaker Verification with Guided Diffusion Model

Yayun He, Zuheng Kang, Jianzong Wang +2

Speaker verification (SV) performance deteriorates as utterances become shorter. To this end, we propose a new architecture called VoiceExtender which provides a promising solution…

cs.SD20231 cited

SVVAD: Personal Voice Activity Detection for Speaker Verification

Zuheng Kang, Jianzong Wang, Junqing Peng +1

Voice activity detection (VAD) improves the performance of speaker verification (SV) by preserving speech segments and attenuating the effects of non-speech. However, this scheme i…

cs.SD20231 cited

Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound Classification

Zuheng Kang, Yayun He, Jianzong Wang +3

Data-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promisin…

cs.SD20221 cited

SpeechEQ: Speech Emotion Recognition based on Multi-scale Unified Datasets and Multitask Learning

Zuheng Kang, Junqing Peng, Jianzong Wang +1

Speech emotion recognition (SER) has many challenges, but one of the main challenges is that each framework does not have a unified standard. In this paper, we propose SpeechEQ, a…