1 paper
Haoyuan Wu, Xinyun Zhang, Peng Xu +3
Vision-Language models (VLMs) pre-trained on large corpora have demonstrated notable success across a range of downstream tasks. In light of the rapidly increasing size of pre-trai…