5 papers
Scaling Spoken Language Models with Syllabic Speech Tokenization
Nicholas Lee, Cheol Jun Cho, Alan W Black +1
Spoken language models (SLMs) typically discretize speech into high-frame-rate tokens extracted from SSL speech models. As the most successful LMs are based on the Transformer arch…
Sylber 2.0: A Universal Syllable Embedding
Cheol Jun Cho, Nicholas Lee, Alan W Black +1
Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolut…
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
Cheol Jun Cho, Abdelrahman Mohamed, Shang-Wen Li +2
Data-driven unit discovery in self-supervised learning (SSL) of speech has embarked on a new era of spoken language processing. Yet, the discovered units often remain in phonetic s…
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
Cheol Jun Cho, Nicholas Lee, Akshat Gupta +4
Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such str…
Coding Speech through Vocal Tract Kinematics
Cheol Jun Cho, Peter Wu, Tejas S. Prabhune +2
Vocal tract articulation is a natural, grounded control space of speech production. The spatiotemporal coordination of articulators combined with the vocal source shapes intelligib…