Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
Yiming Jia, Jiachen Li, Xiang Yue +4
Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-q…
cs.CV2024
LIME: Less Is More for MLLM Evaluation
King Zhu, Qianbo Zang, Shian Jia +18
Multimodal Large Language Models (MLLMs) are evaluated on various benchmarks, such as image captioning, visual question answering, and reasoning. However, many of these benchmarks…