2 papers
cs.CL2026
Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbá
Opeyemi Osakuade, Simon King
Discrete speech units (DSUs) are derived by quantising representations from models trained using self-supervised learning (SSL). They are a popular representation for a wide variet…
cs.CL2024
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
Opeyemi Osakuade, Simon King
Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models, are widely used, especially where there are limited data for the downstream task…