3 papers
cs.IR2026
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Hugo Malard, Michel Olvera, Sanjeel Parekh +3
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at scale. The task, however, remains challengi…
eess.AS2024
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
Hugo Malard, Michel Olvera, Stephane Lathuiliere +1
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the spe…
eess.AS2024
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
Hugo Malard, Michel Olvera, Stéphane Lathuiliere +1
Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In th…