6 citations · 6 across the 1 of their papers we have counts for
1 paper
Yue Zhou, Zhihang Zhong, Xue Yang
Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and…