12 papers
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
Ryo Fukuda, Atsushi Ando, Hiroki Kanagawa +4
Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. How…
SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization
Petr Pálka, Jiangyu Han, Prachi Singh +3
We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…
Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings
Ryo Fukuda, Takatomo Kano, Siddhant Arora +7
We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, tu…
Tight Boundary Prediction in Speaker Diarization Using Causal-Anticausal Consistency
Shota Horiguchi, Marc Delcroix, Naohiro Tawara +2
Multi-talker conversational automatic speech recognition data are often used to train speaker diarization models. Because such data prioritize semantic continuity, pauses and bound…
Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics
Naohiro Tawara, Samuele Cornell, Alexander Polok +3
Conversational automatic speech recognition remains challenging due to overlapping speech, far-field noise, and varying speaker counts. While recent LLM-based systems perform well…
VBx for End-to-End Neural and Clustering-based Diarization
Petr Pálka, Jiangyu Han, Marc Delcroix +2
We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based…