4 papers
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
Yuanzhi Zhu, Xi Wang, Stéphane Lathuilière +1
One-step generators distilled from Masked Diffusion Models (MDMs) compress multiple sampling steps into a single forward pass, enabling efficient text and image synthesis. However,…
Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
Yasser Benigmim, Subhankar Roy, Khalid Oublal +4
The rise of Artificial Intelligence as a Service (AIaaS) democratizes access to pre-trained models via Application Programming Interfaces (APIs), but also raises a fundamental ques…
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
Hugo Malard, Michel Olvera, Stephane Lathuiliere +1
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the spe…