1 paper
Dongchen Lu, Yuyao Sun, Zilu Zhang +4
Most multimodal large language models (MLLMs) treat visual tokens as "a sequence of text", integrating them with text tokens into a large language model (LLM). However, a great qua…