1 paper
Huan Ma, Yan Zhu, Changqing Zhang +5
Vision-language foundation models have exhibited remarkable success across a multitude of downstream tasks due to their scalability on extensive image-text paired data. However, th…