From the 1 of 15 linked papers with an AI index.
15 papers
Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
Stephen McIntosh, Reuben Smit, Daisuke Saito +2
The paper explores using dynamic time warping on self‑supervised WavLM speech representations to automatically score phonetic accuracy, rhythm, and intonation of L2 English and Jap…
Phone Segmentation and Recognition through Phonological Activation Mapping
Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh +8
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phonetic structure is already latent in the re…
Speech Playground: An Interactive Tool for Speech Analysis and Comparison
Stephen McIntosh, Daisuke Saito, Nobuaki Minematsu
This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can be cumbersome to integrate…
SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space
Tomoya Tanabu, Hiroshi Nishijima, Daisuke Saito +1
We introduce SSL-GMMVC, an interpretable voice conversion method in self-supervised speech space. The method models paired source-target features with a Gaussian mixture model and…
Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference
Kentaro Onda, Satoru Fukayama, Daisuke Saito +1
Discrete speech tokens obtained from self-supervised learning (SSL) models provide efficient data compression while maintaining strong performance, and have been widely used as int…
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
Zhijie Huang, Stephen McIntosh, Daisuke Saito +1
A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-lingu…