6 papers
DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification
Qing Wang, Jixun Yao, Zhaokai Sun +3
Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we ai…
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
Pengcheng Guo, Xuankai Chang, Hang Lv +2
Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. Ho…
Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
Shuiyun Liu, Yuxiang Kong, Pengcheng Guo +4
Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses signifi…
MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
Jixun Yao, Qing Wang, Pengcheng Guo +4
Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information…
Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix
Jixun Yao, Qing Wang, Pengcheng Guo +2
Speaker anonymization is an effective privacy protection solution that aims to conceal the speaker's identity while preserving the naturalness and distinctiveness of the original s…
Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder
He Wang, Pengcheng Guo, Xucheng Wan +2
Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use…