1 paper · 1 filter
Jiachen Jiang, Jinxin Zhou, Bo Peng +2
Achieving better alignment between vision embeddings and Large Language Models (LLMs) is crucial for enhancing the abilities of Multimodal LLMs (MLLMs), particularly for recent mod…