3 papers
eess.AS2025
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
Herman Kamper, Benjamin van Niekerk, Julian Zaïdi +1
We introduce LinearVC, a simple voice conversion method that sheds light on the structure of self-supervised representations. First, we show that simple linear transformations of s…
eess.AS2024
Spoken-Term Discovery using Discrete Speech Units
Benjamin van Niekerk, Julian Zaïdi, Marc-André Carbonneau +1
Discovering a lexicon from unlabeled audio is a longstanding challenge for zero-resource speech processing. One approach is to search for frequently occurring patterns in speech. W…
cs.SD2023
EDMSound: Spectrogram Based Diffusion Models for Efficient and High-Quality Audio Synthesis
Ge Zhu, Yutong Wen, Marc-André Carbonneau +1
Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. Thi…