Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
Melody-Lyrics Matching with Contrastive Alignment Loss
Changhong Wang, Michel Olvera, Gaël Richard
The connection between music and lyrics is far beyond semantic bonds. Conceptual pairs in the two modalities such as rhythm and rhyme, note duration and syllabic stress, and struct…
eess.AS2025
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
Hugo Malard, Michel Olvera, Stephane Lathuiliere +1
Large-scale pre-trained audio and image models demonstrate an unprecedented degree of generalization, making them suitable for a wide range of applications. Here, we tackle the spe…
eess.AS2024
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
Hugo Malard, Michel Olvera, Stéphane Lathuiliere +1
Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In th…