4 papers · 1 filter
Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces
Kyle Janse van Rensburg, Herman Kamper
Self-supervised speech features encode both content and speaker information. Recent work introduced an SVD-based factorisation that decomposes these features into a shared content…
Recovering the Zipfian Distribution in Unsupervised Term Discovery
Danel Slabbert, Simon Malan, Herman Kamper
Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Z…
Revisiting Lexicon Evaluation in Unsupervised Word Discovery
Simon Malan, Danel Slabbert, Herman Kamper
Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality?…
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
Matthew Baas, Pieter Scholtz, Arnav Mehta +3
Codec-based text-to-speech (TTS) models have shown impressive quality with zero-shot voice cloning abilities. However, they often struggle with more expressive references or comple…