7 papers
Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces
Kyle Janse van Rensburg, Herman Kamper
Self-supervised speech features encode both content and speaker information. Recent work introduced an SVD-based factorisation that decomposes these features into a shared content…
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
Nicol Visser, Simon Malan, Danel Slabbert +1
Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders re…
Connecting Speech to Words through Images
Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1
How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building…
Recovering the Zipfian Distribution in Unsupervised Term Discovery
Danel Slabbert, Simon Malan, Herman Kamper
Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Z…
Revisiting Lexicon Evaluation in Unsupervised Word Discovery
Simon Malan, Danel Slabbert, Herman Kamper
Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality?…
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
Nicol Visser, Herman Kamper
Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect…