1 paper
Fan Yuan, Yuchen Yan, Yifan Jiang +9
Vision language models (VLMs) achieve unified modeling of images and text, enabling them to accomplish complex real-world tasks through perception, planning, and reasoning. Among t…