1 paper · 1 filter
Timothy Ossowski, Ming Jiang, Junjie Hu
Vision-language models such as CLIP have shown impressive capabilities in encoding texts and images into aligned embeddings, enabling the retrieval of multimodal data in a shared e…