6 papers
Tinted Frames: Question Framing Blinds Vision-Language Models
Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher +2
Vision-Language Models (VLMs) have been shown to be blind, often underutilizing their visual inputs even on tasks that require visual reasoning. In this work, we demonstrate that V…
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
Jiayun Luo, Wan-Cyuan Fan, Lyuyang Wang +4
Large Vision Language Models (LVLMs) have recently emerged as powerful architectures capable of understanding and reasoning over both visual and textual information. These models t…
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan +3
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA),…
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
Wan-Cyuan Fan, Yen-Chun Chen, Mengchen Liu +3
Recent methods for customizing Large Vision Language Models (LVLMs) for domain-specific tasks have shown promising results in scientific chart comprehension. However, existing appr…
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities
Shivam Chandhok, Wan-Cyuan Fan, Vered Shwartz +2
Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable,…
MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
Wan-Cyuan Fan, Tanzila Rahman, Leonid Sigal
With advances in foundational and vision-language models, and effective fine-tuning techniques, a large number of both general and special-purpose models have been developed for a…