1 paper
Lingxiao Zhao, Haoran Zhou, Yuezhi Che +1
Multimodal large language models (MLLMs) extend LLMs with visual understanding through a three-stage pipeline: multimodal preprocessing, vision encoding, and LLM inference. While t…