collaborators

6 papers

cs.SD2025

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification

Qing Wang, Jixun Yao, Zhaokai Sun +3

Being a form of biometric identification, the security of the speaker identification (SID) system is of utmost importance. To better understand the robustness of SID systems, we ai…

eess.AS2024

SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR

Pengcheng Guo, Xuankai Chang, Hang Lv +2

Benefiting from massive and diverse data sources, speech foundation models exhibit strong generalization and knowledge transfer capabilities to a wide range of downstream tasks. Ho…

cs.SD2024

Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge

Shuiyun Liu, Yuxiang Kong, Pengcheng Guo +4

Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses signifi…

eess.AS2024

MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement

Jixun Yao, Qing Wang, Pengcheng Guo +4

Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information…

eess.AS2024

Distinctive and Natural Speaker Anonymization via Singular Value Transformation-assisted Matrix

Jixun Yao, Qing Wang, Pengcheng Guo +2

Speaker anonymization is an effective privacy protection solution that aims to conceal the speaker's identity while preserving the naturalness and distinctiveness of the original s…

cs.CV2024

Enhancing Lip Reading with Multi-Scale Video and Multi-Encoder

He Wang, Pengcheng Guo, Xucheng Wan +2

Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use…