1 paper
Xiangyang Wu, Liu Liu, Baosheng Yu +2
Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within…