1 paper
Kevin Robbins, Xiaotong Liu, Yu Wu +4
Vision-Language Models like CLIP create aligned embedding spaces for text and images, making it possible for anyone to build a visual classifier by simply naming the classes they w…