Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
Kuei-Chun Kao, Hsu Tzu-Yin, Yunqi Hong +2
Recently, Multimodal Large Language Models (MLLMs) encounter two key issues in multi-image contexts: (1) a lack of fine-grained perception across disparate images, and (2) a dimini…
cs.CV2025
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
Andrew Bai, Justin Cui, Ruochen Wang +1
Vision-language instruction tuning achieves two main purposes: learning visual concepts and learning visual skills. In this paper, we found that vision-language benchmarks fall int…