activity
20242026
collaborators

12 papers

cs.SD2026

InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective

Samir Sadok, Xavier Alameda-Pineda

Self-supervised learning (SSL) models, such as Wav2Vec2, HuBERT, and WavLM, have become foundational across a wide range of speech and audio tasks. Despite their success, understan…

cs.SD2026

Diffusion-based Frameworks for Unsupervised Speech Enhancement

Jean-Eudes Ayilo, Mostafa Sadeghi, Romain Serizel +1

This paper addresses unsupervised diffusion-based single-channel speech enhancement (SE). Prior work in this direction combines a score-based diffusion model trained on clean speec…

cs.SD2026

Modeling strategies for speech enhancement in the latent space of a neural audio codec

Sofiene Kammoun, Xavier Alameda-Pineda, Simon Leglaive

Neural audio codecs (NACs) provide compact latent speech representations in the form of sequences of continuous vectors or discrete tokens. In this work, we investigate how these t…

cs.SD2026

The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

Samir Sadok, Laurent Girin, Xavier Alameda-Pineda

Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space. As a result,…

cs.SD2026

Residual Tokens Enhance Masked Autoencoders for Speech Modeling

Samir Sadok, Stéphane Lathuilière, Xavier Alameda-Pineda

Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce…

cs.AI2026

OpenSocInt: A Multi-modal Training Environment for Human-Aware Social Navigation

Victor Sanchez, Chris Reinke, Ahamed Mohamed +1

In this paper, we introduce OpenSocInt, an open-source software package providing a simulator for multi-modal social interactions and a modular architecture to train social agents.…