Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
Haitong Sun, Stephen McIntosh, Kwanghee Choi +3
Speech representations from self-supervised speech models (S3Ms) are known to be sensitive to phonemic contrasts, but their sensitivity to prosodic contrasts has not been directly…
cs.CL2026
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
Zhijie Huang, Stephen McIntosh, Daisuke Saito +1
A good language model starts with a good tokenizer. Tokenization is especially important for speech modeling, which must handle continuous signals that mix linguistic and non-lingu…
cs.CL2024
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
Joonyong Park, Daisuke Saito, Nobuaki Minematsu
We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations…