1 paper
Tingwei Zhang, Rishi Jha, Eugene Bagdasaryan +1
Multi-modal embeddings encode texts, images, thermal images, sounds, and videos into a single embedding space, aligning representations across different modalities (e.g., associate…