16 citations · 16 across the 1 of their papers we have counts for
1 paper
Omiros Pantazis, Gabriel Brostow, Kate Jones +1
Vision-language models such as CLIP are pretrained on large volumes of internet sourced image and text pairs, and have been shown to sometimes exhibit impressive zero- and low-shot…