1 paper · 1 filter
Morris Alper, Hadar Averbuch-Elor
While recent vision-and-language models (VLMs) like CLIP are a powerful tool for analyzing text and images in a shared semantic space, they do not explicitly model the hierarchical…