21 citations · 21 across the 1 of their papers we have counts for
1 paper
Federico Bianchi, Giuseppe Attanasio, Raphael Pisoni +3
CLIP (Contrastive Language-Image Pre-training) is a very recent multi-modal model that jointly learns representations of images and texts. The model is trained on a massive amount…