1 paper
Mingjia Shi, Shuo Wang, Xiaobo Wang +7
Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requiring full-parameter updates.…