papers

Publications (47)

cs.IR2026

Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

Hugo Malard, Michel Olvera, Sanjeel Parekh +3

Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…

eess.AS2023

Fine-tuning Strategies for Faster Inference using Speech Self-Supervised Models: A Comparative Study

Salah Zaiem, Robin Algayres, Titouan Parcollet +2

Self-supervised learning (SSL) has allowed substantial progress in Automatic Speech Recognition (ASR) performance in low-resource settings. In this context, it has been demonstrate…

cs.SD2024

A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification

Michel Olvera, Paraskevas Stamatiadis, Slim Essid

Audio-text models trained via contrastive learning offer a practical approach to perform audio classification through natural language prompts, such as "this is a sound of" followe…

eess.AS2024

Online speaker diarization of meetings guided by speech separation

Elio Gruttadauria, Mathieu Fontaine, Slim Essid

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Al…

cs.LG2018

Structured Output Learning with Abstention: Application to Accurate Opinion Prediction

Alexandre Garcia, Slim Essid, Chloé Clavel +1

Motivated by Supervised Opinion Analysis, we propose a novel framework devoted to Structured Output Learning with Abstention (SOLA). The structure prediction model is able to absta…

cs.SD2024

A lightweight dual-stage framework for personalized speech enhancement based on DeepFilterNet2

Thomas Serre, Mathieu Fontaine, Éric Benhaim +2

Isolating the desired speaker's voice amidst multiplespeakers in a noisy acoustic context is a challenging task. Per-sonalized speech enhancement (PSE) endeavours to achievethis by…

cs.CV2018

Weakly Supervised Representation Learning for Unsynchronized Audio-Visual Events

Sanjeel Parekh, Slim Essid, Alexey Ozerov +3

Audio-visual representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel…

cs.CV2024

Collaborating Foundation Models for Domain Generalized Semantic Segmentation

Yasser Benigmim, Subhankar Roy, Slim Essid +2

Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGS…

cs.LG2026

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

Victor Letzelter, Hugo Malard, Mathieu Fontaine +4

We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference…

cs.CV2025

Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation

Yasser Benigmim, Subhankar Roy, Khalid Oublal +4

The rise of Artificial Intelligence as a Service (AIaaS) democratizes access to pre-trained models via Application Programming Interfaces (APIs), but also raises a fundamental ques…

cs.HC2020

On-the-fly Detection of User Engagement Decrease in Spontaneous Human-Robot Interaction, International Journal of Social Robotics, 2019

Atef Ben Youssef, Giovanna Varni, Slim Essid +1

In this paper, we consider the detection of a decrease of engagement by users spontaneously interacting with a socially assistive robot in a public space. We first describe the UE-…

cs.CL2019

From the Token to the Review: A Hierarchical Multimodal approach to Opinion Mining

Alexandre Garcia, Pierre Colombo, Slim Essid +2

The task of predicting fine grained user opinion based on spontaneous spoken language is a key problem arising in the development of Computational Agents as well as in the developm…

eess.SP2021

Attention-based distributed speech enhancement for unconstrained microphone arrays with varying number of nodes

Nicolas Furnon, Romain Serizel, Slim Essid +1

Speech enhancement promises higher efficiency in ad-hoc microphone arrays than in constrained microphone arrays thanks to the wide spatial coverage of the devices in the acoustic s…

eess.AS2023

Automatic Data Augmentation for Domain Adapted Fine-Tuning of Self-Supervised Speech Representations

Salah Zaiem, Titouan Parcollet, Slim Essid

Self-Supervised Learning (SSL) has allowed leveraging large amounts of unlabeled speech data to improve the performance of speech recognition models even with small annotated datas…

cs.CV2023

One-shot Unsupervised Domain Adaptation with Personalized Diffusion Models

Yasser Benigmim, Subhankar Roy, Slim Essid +2

Adapting a segmentation model from a labeled source domain to a target domain, where a single unlabeled datum is available, is one the most challenging problems in domain adaptatio…

cs.AI2026

S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models

Mohammed Ali El Adlouni, Aurian Quelennec, Pierre Chouteau +2

General audio foundation models have recently achieved remarkable progress, enabling strong performance across diverse tasks. However, state-of-the-art models remain extremely larg…

eess.AS2022

Automatic Data Augmentation Selection and Parametrization in Contrastive Self-Supervised Speech Representation Learning

Salah Zaiem, Titouan Parcollet, Slim Essid

Contrastive learning enables learning useful audio and speech representations without ground-truth labels by maximizing the similarity between latent representations of similar sig…

cs.SD2025

MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning

Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters +1

Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods h…

eess.AS2025

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization

Hugo Malard, Michel Olvera, Stephane Lathuiliere +1

Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the spe…

cs.SD2025

Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping

Clémentine Berger, Roland Badeau, Slim Essid

People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise's frequency components due…

cs.SD2025

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters +1

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned l…

eess.SP2020

DNN-based mask estimation for distributed speech enhancement in spatially unconstrained microphone arrays

Nicolas Furnon, Romain Serizel, Irina Illina +1

Deep neural network (DNN)-based speech enhancement algorithms in microphone arrays have now proven to be efficient solutions to speech understanding and speech recognition in noisy…

cs.SD2024

Multiple Choice Learning for Efficient Speech Separation with Many Speakers

David Perera, François Derrida, Théo Mariotte +2

Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated…

cs.CL2018

Opinion Dynamics Modeling for Movie Review Transcripts Classification with Hidden Conditional Random Fields

Valentin Barriere, Chloé Clavel, Slim Essid

In this paper, the main goal is to detect a movie reviewer's opinion using hidden conditional random fields. This model allows us to capture the dynamics of the reviewer's opinion…

cs.CV2018

Identify, locate and separate: Audio-visual object extraction in large video collections using weak supervision

Sanjeel Parekh, Alexey Ozerov, Slim Essid +3

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object…

eess.AS2024

Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads

Salah Zaiem, Youcef Kemiche, Titouan Parcollet +2

Self-supervised learning (SSL) leverages large datasets of unlabeled speech to reach impressive performance with reduced amounts of annotated data. The high number of proposed appr…

cs.SD2024

SALT: Standardized Audio event Label Taxonomy

Paraskevas Stamatiadis, Michel Olvera, Slim Essid

Machine listening systems often rely on fixed taxonomies to organize and label audio data, key for training and evaluating deep neural networks (DNNs) and other supervised algorith…

cs.SD2025

Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking

Antonin Gagnere, Slim Essid, Geoffroy Peeters

Ambiguities in data and problem constraints can lead to diverse, equally plausible outcomes for a machine learning task. In beat and downbeat tracking, for instance, different list…

eess.AS2024

Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations

Salah Zaiem, Titouan Parcollet, Slim Essid

Despite being trained on massive and diverse datasets, speech self-supervised encoders are generally used for downstream purposes as mere frozen feature extractors or model initial…

eess.AS2022

Pretext Tasks selection for multitask self-supervised speech representation learning

Salah Zaiem, Titouan Parcollet, Slim Essid +1

Through solving pretext tasks, self-supervised learning leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream tas…

eess.AS2023

Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?

Salah Zaiem, Youcef Kemiche, Titouan Parcollet +2

Self-supervised learning (SSL) has recently allowed leveraging large datasets of unlabeled speech signals to reach impressive performance on speech tasks using only small amounts o…

eess.SP2021

Distributed speech separation in spatially unconstrained microphone arrays

Nicolas Furnon, Romain Serizel, Irina Illina +1

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current st…

eess.AS2024

An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment

Hugo Malard, Michel Olvera, Stéphane Lathuiliere +1

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In th…

cs.SD2026

Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement

Thomas Serre, Mathieu Fontaine, Éric Benhaim +1

Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorp…

cs.LG2025

Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing

David Perera, Victor Letzelter, Théo Mariotte +4

We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of…

cs.SD2023

On the choice of the optimal temporal support for audio classification with Pre-trained embeddings

Aurian Quelennec, Michel Olvera, Geoffroy Peeters +1

Current state-of-the-art audio analysis systems rely on pre-trained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of ta…

eess.AS2021

Conditional independence for pretext task selection in Self-supervised speech representation learning

Salah Zaiem, Titouan Parcollet, Slim Essid

Through solving pretext tasks, self-supervised learning (SSL) leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstre…

cs.SD2026

TinyMU: A Compact Audio-Language Model for Music Understanding

Xiquan Li, Aurian Quelennec, Slim Essid

Music understanding and reasoning are central challenges in the Music Information Research field, with applications ranging from retrieval and recommendation to music agents and vi…

cs.MM2021

A multimodal movie review corpus for fine-grained opinion mining

Alexandre Garcia, Slim Essid, Florence d'Alché-Buc +1

In this paper, we introduce a set of opinion annotations for the POM movie review dataset, composed of 1000 videos. The annotation campaign is motivated by the development of a hie…

stat.ML2023

Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis

Victor Letzelter, Mathieu Fontaine, Mickaël Chen +3

We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may b…

cs.LG2024

Winner-takes-all learners are geometry-aware conditional density estimators

Victor Letzelter, David Perera, Cédric Rommel +4

Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between W…

q-bio.NC2018

EEG-based Inter-Subject Correlation Schemes in a Stimuli-Shared Framework: Interplay with Valence and Arousal

Ayoub Hajlaoui, Mohamed Chetouani, Slim Essid

Affective computing is confronted to high inter-subject variability, in both emotional and physiological responses to a given stimulus. In a stimuli-shared framework, that is to sa…

eess.AS2025

IS : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering

Clémentine Berger, Paraskevas Stamatiadis, Roland Badeau +1

We are interested in audio systems capable of performing a differentiated processing of stationary backgrounds and isolated acoustic events within an acoustic scene, whether for ap…

eess.AS2023

SAMbA: Speech enhancement with Asynchronous ad-hoc Microphone Arrays

Nicolas Furnon, Romain Serizel, Slim Essid +1

Speech enhancement in ad-hoc microphone arrays is often hindered by the asynchronization of the devices composing the microphone array. Asynchronization comes from sampling time of…

cs.LG2025

O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization

Elio Gruttadauria, Mathieu Fontaine, Jonathan Le Roux +1

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we…

cs.SD2020

DNN-Based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays

Nicolas Furnon, Romain Serizel, Irina Illina +1

Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions to the real-world. Distributed sensor arrays that…

eess.AS2024

A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning

Antonin Gagnere, Geoffroy Peeters, Slim Essid

In this paper, we propose a novel Self-Supervised-Learning scheme to train rhythm analysis systems and instantiate it for few-shot beat tracking. Taking inspiration from the Contra…