5 papers
StyleStream: Real-Time Zero-Shot Voice Style Conversion
Yisi Liu, Nicholas Lee, Gopala Anumanchipalli
Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguisti…
Scaling Spoken Language Models with Syllabic Speech Tokenization
Nicholas Lee, Cheol Jun Cho, Alan W Black +1
Spoken language models (SLMs) typically discretize speech into high-frame-rate tokens extracted from SSL speech models. As the most successful LMs are based on the Transformer arch…
Sylber 2.0: A Universal Syllable Embedding
Cheol Jun Cho, Nicholas Lee, Alan W Black +1
Scaling spoken language modeling requires speech tokens that are both efficient and universal. Recent work has proposed syllables as promising speech tokens at low temporal resolut…
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
Immanuel Abdi, Akshat Gupta, Micah Mok +3
One of the biggest missing capabilities in current AI systems is the ability to learn continuously after deployment. Implementing such continually learning systems have several cha…
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
Cheol Jun Cho, Nicholas Lee, Akshat Gupta +4
Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such str…