1 paper · 1 filter
Qiao Liang, Yanjiang Liu, Weixiang Zhou +7
Does the prior knowledge of the vision encoder constrain the capability boundary of Multi-modal Large Language Models (MLLMs)? While most existing research treats MLLMs as unified…