1 paper
Parthasaarathy Sudarsanam, Irene Martín-Morató, Tuomas Virtanen
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive trainin…