3 papers
cs.CV2026
VGR: Visual Grounded Reasoning
Jiacong Wang, Zijian Kang, Haochen Wang +8
In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias…
cs.CV2024
Unveiling the Tapestry of Consistency in Large Vision-Language Models
Yuan Zhang, Fei Xiao, Tao Huang +7
Large vision-language models (LVLMs) have recently achieved rapid progress, exhibiting great perception and reasoning abilities concerning visual information. However, when faced w…
cs.CV2024
Benchmarking and Improving Detail Image Caption
Hongyuan Dong, Jiawen Li, Bohong Wu +3
Image captioning has long been regarded as a fundamental task in visual understanding. Recently, however, few large vision-language model (LVLM) research discusses model's image ca…