1 citations · 1 across the 4 of their papers we have counts for
4 papers
Vision Language Models Cannot Reason About Physical Transformation
Dezhi Luo, Yijiang Li, Maijunxian Wang +7
Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they…
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
Hung Huy Nguyen, Pooyan Rahmanzadehgervi, Long Mai +1
Detecting object-level changes between two images across possibly different views is a core task in many applications that involve visual inspection or camera surveillance. Existin…
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
Pooyan Rahmanzadehgervi, Hung Huy Nguyen, Rosanne Liu +2
Multi-head self-attention (MHSA) is a key component of Transformers, a widely popular architecture in both language and vision. Multiple heads intuitively enable different parallel…
Vision language models are blind: Failing to translate detailed visual features into words
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1
While large language models with vision capabilities (VLMs), e.g., GPT-4o and Gemini 1.5 Pro, score high on many vision-understanding benchmarks, they are still struggling with low…