works on

From the 1 of 15 linked papers with an AI index.

activity
20242026
collaborators

15 papers

cs.CL2026

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

Stephen McIntosh, Reuben Smit, Daisuke Saito +2

The paper explores using dynamic time warping on self‑supervised WavLM speech representations to automatically score phonetic accuracy, rhythm, and intonation of L2 English and Jap…

eess.AS2026

Phone Segmentation and Recognition through Phonological Activation Mapping

Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8

Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the re…

cs.CL2026

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

Stephen McIntosh, Daisuke Saito, Nobuaki Minematsu

This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can be cumbersome to integrate…

eess.AS2026

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

Tomoya Tanabu, Hiroshi Nishijima, Daisuke Saito +1

We introduce SSL-GMMVC, an interpretable voice conversion method in self-supervised speech space. The method models paired source-target features with a Gaussian mixture model and…

cs.SD2026

Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

Kentaro Onda, Satoru Fukayama, Daisuke Saito +1

Discrete speech tokens obtained from self-supervised learning (SSL) models provide efficient data compression while maintaining strong performance, and have been widely used as int…

cs.CL2026

Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

Zhijie Huang, Stephen McIntosh, Daisuke Saito +1

A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-lingu…