collaborators

6 papers

cs.CL2026

ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling

Nicol Visser, Simon Malan, Danel Slabbert +1

Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders re…

eess.AS2026

Recovering the Zipfian Distribution in Unsupervised Term Discovery

Danel Slabbert, Simon Malan, Herman Kamper

Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Z…

eess.AS2026

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

Simon Malan, Danel Slabbert, Herman Kamper

Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality?…

eess.AS2026

Unsupervised lexicon learning from speech is limited by representations rather than clustering

Danel Slabbert, Simon Malan, Herman Kamper

Zero-resource word segmentation and clustering systems aim to tokenise speech into word-like units without access to text labels. Despite progress, the induced lexicons are still f…

eess.AS2025

Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?

Simon Malan, Benjamin van Niekerk, Herman Kamper

We investigate the problem of segmenting unlabeled speech into word-like units and clustering these to create a lexicon. Prior work can be categorized into two frameworks. Bottom-u…

eess.AS2025

Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming

Simon Malan, Benjamin van Niekerk, Herman Kamper

We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model couple…