1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Zhuokun Chen, Jinwu Hu, Zeshuai Deng +3
Multimodal LLMs (MLLMs) equip language models with visual capabilities by aligning vision encoders with language models. Existing methods to enhance the visual perception of MLLMs…