1 paper
Zheng Ma, Mianzhi Pan, Wenhan Wu +4
Vision-language models (VLMs) have shown impressive performance in substantial downstream multi-modal tasks. However, only comparing the fine-tuned performance on downstream tasks…