15 papers
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
Victor Letzelter, Hugo Malard, Mathieu Fontaine +4
We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference…
S-SONDO: Self-Supervised Knowledge Distillation for General Audio Foundation Models
Mohammed Ali El Adlouni, Aurian Quelennec, Pierre Chouteau +2
General audio foundation models have recently achieved remarkable progress, enabling strong performance across diverse tasks. However, state-of-the-art models remain extremely larg…
TinyMU: A Compact Audio-Language Model for Music Understanding
Xiquan Li, Aurian Quelennec, Slim Essid
Music understanding and reasoning are central challenges in the Music Information Research field, with applications ranging from retrieval and recommendation to music agents and vi…
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
Thomas Serre, Mathieu Fontaine, Ãric Benhaim +1
Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorp…
O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
Elio Gruttadauria, Mathieu Fontaine, Jonathan Le Roux +1
We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we…