Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
Nicol Visser, Simon Malan, Danel Slabbert +1
Pure speech language models aim to learn language directly from raw audio without textual resources. A key challenge is that discrete tokens from self-supervised speech encoders re…
cs.CL2026
Connecting Speech to Words through Images
Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1
How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building…
cs.CL2025
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
Nicol Visser, Herman Kamper
Spoken language models (SLMs) operate on acoustic units obtained by discretizing self-supervised speech representations. Although the characteristics of these units directly affect…