1 citations · 1 across the 1 of their papers we have counts for
1 paper
Antonio D'Orazio, Maria Rosaria Briglia, Donato Crisostomi +3
CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines…