From the 1 of 12 linked papers with an AI index.
12 papers
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…
SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization
Petr Pálka, Jiangyu Han, Prachi Singh +3
We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4
Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
Jiangyu Han, Petr Pálka, Marc Delcroix +4
Self-supervised learning (SSL) models such as WavLM have substantially advanced speaker diarization by providing rich contextual speech representations. However, the high computati…
Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
Junyi Peng, Lin Zhang, Jiangyu Han +5
Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on…
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
Yoshiki Masuyama, Kohei Saijo, Francesco Paissan +6
Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configur…