14 papers
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…
Multiple Choice Learning of Low-Rank Adapters for Language Modeling
Victor Letzelter, Hugo Malard, Mathieu Fontaine +4
We propose LoRA-MCL, a training scheme that extends next-token prediction in language models with a method designed to decode diverse, plausible sentence continuations at inference…
Taylor-SWFT: fast discrete Statistical Wave Field Theory using Taylor expansion for late reverberation Work under review
Marius Rodrigues, Louis Lalay, Roland Badeau +2
Dynamic room acoustic simulation aims to render the acoustic effects of an environment in real time while accounting for potentially moving sources and receivers. In this context,…
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
Louis Bahrman, Marius Rodrigues, Mathieu Fontaine +1
This paper explores the outcome of training state-of-the-art dereverberation models with supervision settings ranging from weakly-supervised to virtually unsupervised, relying sole…
S-PRESSO: Ultra Low Bitrate Sound Effect Compression With Diffusion Autoencoders And Offline Quantization
Zineb Lahrichi, Gaëtan Hadjeres, Gaël Richard +1
Neural audio compression models have recently achieved extreme compression rates, enabling efficient latent generative modeling. Conversely, latent generative models have been appl…
GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
Teysir Baoueb, Xiaoyu Bie, Mathieu Fontaine +1
Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stabili…