20 citations · 20 across the 4 of their papers we have counts for
5 papers · 1 filter
SO-Bench: A Structural Output Evaluation of Multimodal LLMs
Di Feng, Kaixin Ma, Feng Nan +9
Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to predefined data schem…
Law of Vision Representation in MLLMs
Shijia Yang, Bohan Zhai, Quanzeng You +3
We present the "Law of Vision Representation" in multimodal large language models (MLLMs). It reveals a strong correlation between the combination of cross-modal alignment, corresp…
InfiMM-HD: A Leap Forward in High-Resolution Multimodal Understanding
Haogeng Liu, Quanzeng You, Xiaotian Han +7
Multimodal Large Language Models (MLLMs) have experienced significant advancements recently. Nevertheless, challenges persist in the accurate recognition and comprehension of intri…
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
Xiaotian Han, Yiqi Wang, Bohan Zhai +2
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. Visual instruction fine-tuning (IFT) is a vital process for aligning M…
InfiMM-Eval: Complex Open-Ended Reasoning Evaluation For Multi-Modal Large Language Models
Xiaotian Han, Quanzeng You, Yongfei Liu +9
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. These models not only excel in traditional vision-language tasks but a…