activity
20242026
collaborators

16 papers

eess.AS2026

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

Ye-Xin Lu, Xin Wang, Yang Ai +3

Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or enta…

cs.CV2026

SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment

Sahibzada Adil Shahzad, Ammarah Hashmi, Junichi Yamagishi +5

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on cur…

cs.SD2026

Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction

Yun Liu, Xuechen Liu, Xiaoxiao Miao +1

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due…

cs.CL2025

AfriHuBERT: A self-supervised speech representation model for African languages

Jesujoba O. Alabi, Xuechen Liu, Dietrich Klakow +1

In this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages. While mHuBERT-147 covered 16 African l…

eess.AS2025

The First VoicePrivacy Attacker Challenge

Natalia Tomashenko, Xiaoxiao Miao, Emmanuel Vincent +1

The First VoicePrivacy Attacker Challenge is an ICASSP 2025 SP Grand Challenge which focuses on evaluating attacker systems against a set of voice anonymization systems submitted t…

cs.SD2024

Explaining Speaker and Spoof Embeddings via Probing

Xuechen Liu, Junichi Yamagishi, Md Sahidullah +1

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as…