9 papers
InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective
Samir Sadok, Xavier Alameda-Pineda
Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understan…
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Artem Ploujnikov, Francesco Verdini, Samir Sadok +1
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabilities into Large Language Models (LLMs). How…
Data-Driven Decoding of Russell's Circumplex Model of Affect
Amdjed Belaref, Samir Sadok, Zineb Noumir +1
Affective computing increasingly relies on deep learning to represent emotions, yet latent spaces often remain opaque, high-dimensional black boxes. This paper investigates whether…
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation
Pedro Correa, Olivier Perrotin, Samir Sadok +2
The choice of speech representation is critical in speech-driven 3D facial animation. Representations differ in what they encode: SSL features emphasize segmental and semantic cues…
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
Samir Sadok, Laurent Girin, Xavier Alameda-Pineda
Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result,…
Residual Tokens Enhance Masked Autoencoders for Speech Modeling
Samir Sadok, Stéphane Lathuilière, Xavier Alameda-Pineda
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce…