1 paper
Zilun Zhang, Cuifeng Shen, Yuan Shen +4
Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224),…