works on

From the 1 of 14 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring

Stephen McIntosh, Reuben Smit, Daisuke Saito +2

The paper explores using dynamic time warping on self‑supervised WavLM speech representations to automatically score phonetic accuracy, rhythm, and intonation of L2 English and Jap…

cs.CL2026

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

Stephen McIntosh, Daisuke Saito, Nobuaki Minematsu

This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can be cumbersome to integrate…

cs.CL2026

Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

Zhijie Huang, Stephen McIntosh, Daisuke Saito +1

A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-lingu…

cs.CL2026

Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations

Haitong Sun, Stephen McIntosh, Kwanghee Choi +3

Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly…

cs.CL2024

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model

Joonyong Park, Daisuke Saito, Nobuaki Minematsu

We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations…