1 paper
Jingyuan Qi, Zhiyang Xu, Rulin Shao +5
Current vision-language models (VLMs) still exhibit inferior performance on knowledge-intensive tasks, primarily due to the challenge of accurately encoding all the associations be…