works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CL2026

M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models

Ryo Fukuda, Atsushi Ando, Hiroki Kanagawa +4

Full-duplex spoken dialogue systems (FDSDSs) can listen while speaking, enabling natural behaviors such as smooth turn-taking, backchannel handling, and user barge-in handling. How…

eess.AS2026

SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Shuai Wang, Zihan Qian, Ke Zhang +9

The paper presents the REAL‑TSE Challenge, a benchmark for extracting a target speaker’s voice from real conversational recordings in Mandarin and English, with both online low‑lat…

eess.AS2026

SphereVBx: Spherical Variational Bayes Clustering for Simplified EEND-VC Diarization

Petr Pálka, Jiangyu Han, Prachi Singh +3

We propose SphereVBx, a Bayesian clustering framework for hyperspherical embeddings based on Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA). The method follows the…

cs.CL2026

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

Ryo Fukuda, Takatomo Kano, Siddhant Arora +7

We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, tu…

eess.AS2026

Non-Autoregressive Minimum Bayes' Risk Decoding for Fast Speech Recognition

Hiroyuki Deguchi, Takatomo Kano, Katsuki Chousa +1

Non-autoregressive (NAR) decoding generates output tokens in parallel, making speech recognition faster than autoregressive decoding, which generates them sequentially from left to…

eess.AS2026

Generating Training Targets for Real-World Speech Enhancement via Close-to-Distant Microphone Projection

Tomohiro Nakatani, Rintaro Ikeshita, Naoyuki Kamo +2

Training neural networks (NNs) for speech enhancement (SE) in distant speech-capturing scenarios requires paired distorted and clean reference speech signals. While such data are o…