12 papers · 1 filter
Diarization Error Decomposition Under Pause Annotation Ambiguity
Shota Horiguchi, Marc Delcroix, Naohiro Tawara +1
Speaker diarization evaluation is sensitive to ambiguity in pause annotation, which can inflate diarization error rate (DER) or obscure genuine model errors. We show that morpholog…
Over-Tightening-Aware Pseudo-Labeling for Tight-Boundary Speaker Diarization
Shota Horiguchi, Takanori Ashihara, Marc Delcroix +2
Training speaker diarization models on loose labels, such as speech segments with padded boundaries or filled pauses, often results in similarly loose model outputs. To obtain tigh…
Ontology-based Target Sound Extraction
Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai +2
Target sound extraction (TSE) aims to isolate a sound source of interest from a mixture, given a semantic query. Existing TSE systems are conditioned on fixed class representations…
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Shuai Wang, Zihan Qian, Ke Zhang +9
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and…
SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization
Petr Pálka, Jiangyu Han, Prachi Singh +3
We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…
Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition
Hiroyuki Deguchi, Takatomo Kano, Katsuki Chousa +1
Non-autoregressive (NAR) decoding generates output tokens in parallel, making speech recognition faster than autoregressive decoding, which generates them sequentially from left to…