works on

From the 1 of 24 linked papers with an AI index.

activity
20242026
collaborators

24 papers

cs.SD2026

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

Joonyong Park, David M. Chan, Yuki Saito +1

The paper investigates how large audio-language models used as automatic judges for speech evaluation can exploit protocol-level shortcuts—relying on provided labels or reference d…

cs.CL2026

DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling

Wataru Nakata, Yuki Saito, Hiroshi Saruwatari

Full-duplex spoken dialogue models are trained on conversational speech in which each speaker is represented as a separate stream, but existing large-scale public speech corpora ar…

cs.SD2026

DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio

Wataru Nakata, Yuki Saito, Kazuki Yamauchi +2

Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most…

eess.AS2026

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

Dong Yang, Yuki Saito, Wataru Nakata +1

This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and…

eess.AS2026

Fast Multichannel NMF with Block-Diagonal Spatial Covariance Matrices for Efficient Blind Source Separation Using Distributed Microphone Arrays

Hirotaka Nishikori, Nobutaka Ito, Kouei Yamaoka +2

Distributed microphone arrays composed of multiple subarrays enable blind source separation over a wide spatial area. Directly applying fast multichannel nonnegative matrix factori…

eess.AS2026

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

Dong Yang, Yiyi Cai, Haoyu Zhang +2

Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requirin…