5 papers
On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models
Shunsuke Kando, Wataru Nakata, Shinnosuke Takamichi +1
Generative Spoken Language Modeling (GSLM) enables text-free speech modeling by training language models (LMs) using discrete speech representations instead of textual transcriptio…
Analysing the Language of Neural Audio Codecs
Joonyong Park, Shinnosuke Takamichi, David M. Chan +3
This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC…
Do Self-Supervised Speech Models Exhibit the Critical Period Effects in Language Acquisition?
Yurie Koga, Shunsuke Kando, Yusuke Miyao
This paper investigates whether the Critical Period (CP) effects in human language acquisition are observed in self-supervised speech models (S3Ms). CP effects refer to greater dif…
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi
The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While s…
Syntactic Learnability of Echo State Neural Language Models at Scale
Ryo Ueda, Tatsuki Kuribayashi, Shunsuke Kando +1
What is a neural model with minimum architectural complexity that exhibits reasonable language learning capability? To explore such a simple but sufficient neural language model, w…