6 papers · 1 filter
InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective
Samir Sadok, Xavier Alameda-Pineda
Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understan…
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
Samir Sadok, Laurent Girin, Xavier Alameda-Pineda
Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result,…
Residual Tokens Enhance Masked Autoencoders for Speech Modeling
Samir Sadok, Stéphane Lathuilière, Xavier Alameda-Pineda
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce…
A vector quantized masked autoencoder for audiovisual speech emotion recognition
Samir Sadok, Simon Leglaive, Renaud Séguier
An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised mult…
AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder
Samir Sadok, Simon Leglaive, Laurent Girin +2
This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen…
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
Samir Sadok, Simon Leglaive, Laurent Girin +2
In this paper, we present a multimodal and dynamical VAE (MDVAE) applied to unsupervised audio-visual speech representation learning. The latent space is structured to dissociate t…