2 papers
cs.AI2026
Thinking with Visual Grounding
Junkai Zhang, Yihe Deng, Kai-Wei Chang +1
Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces oft…
cs.CV2024
Enhancing Large Vision Language Models with Self-Training on Image Comprehension
Yihe Deng, Pan Lu, Fan Yin +6
Large vision language models (LVLMs) integrate large language models (LLMs) with pre-trained vision encoders, thereby activating the perception capability of the model to understan…