activity
20202026
most citedImproving speaker discrimination of target speech extraction with time-domain SpeakerBeam

12 citations · 16 across the 27 of their papers we have counts for

collaborators

30 papers

eess.AS2026

Diarization Error Decomposition Under Pause Annotation Ambiguity

Shota Horiguchi, Marc Delcroix, Naohiro Tawara +1

Speaker diarization evaluation is sensitive to ambiguity in pause annotation, which can inflate diarization error rate (DER) or obscure genuine model errors. We show that morpholog…

eess.AS2026

Over-Tightening-Aware Pseudo-Labeling for Tight-Boundary Speaker Diarization

Shota Horiguchi, Takanori Ashihara, Marc Delcroix +2

Training speaker diarization models on loose labels, such as speech segments with padded boundaries or filled pauses, often results in similarly loose model outputs. To obtain tigh…

eess.AS2026

Ontology-based Target Sound Extraction

Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai +2

Target sound extraction (TSE) aims to isolate a sound source of interest from a mixture, given a semantic query. Existing TSE systems are conditioned on fixed class representations…

cs.CL2026

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

Ryo Fukuda, Atsushi Ando, Hiroki Kanagawa +4

Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. How…

eess.AS2026

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

Petr Pálka, Jiangyu Han, Prachi Singh +3

We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…

cs.CL2026

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

Ryo Fukuda, Takatomo Kano, Siddhant Arora +7

We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, tu…