1 paper
Jianhong Tu, Zhuohao Ni, Nicholas Crispino +8
We present a novel visual instruction tuning strategy to improve the zero-shot task generalization of multimodal large language models by building a firm text-only knowledge base.…