10 papers
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
Jiadong Wang, Ke Zhang, Xinyuan Qian +3
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Alt…
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3
Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Junjie Li +3
In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during tr…
Leveraging Language Information for Target Language Extraction
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang +2
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory sy…
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
Jiawei Zhang, Tian-Hao Zhang, Jun Wang +3
Controlling the style and characteristics of speech synthesis is crucial for adapting the output to specific contexts and user requirements. Previous Text-to-speech (TTS) works hav…
Interpolating Speaker Identities in Embedding Space for Data Expansion
Tianchi Liu, Ruijie Tao, Qiongqiong Wang +5
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more…