activity
20242026
most citedVisual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs

1 citations · 2 across the 9 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

Causal Attribution via Activation Patching

Amirmohammad Izadi, Mohammadali Banayeeanzade, Alireza Mirrokni +4

Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains ch…

cs.CV2025

Understanding Counting Mechanisms in Large Language and Vision-Language Models

Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari +4

Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and…

cs.CV2025

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari +4

Large vision-language models (LVLMs) show strong performance across multimodal benchmarks but remain limited in structured reasoning and precise grounding. Recent work has demonstr…

cs.CV2025

Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs

Amirmohammad Izadi, Mohammad Ali Banayeeanzade, Fatemeh Askari +4

Despite progress in Large Vision-Language Models (LVLMs), their capacity for visual reasoning is often limited by the binding problem: the failure to reliably associate perceptual…

cs.CV2025

T2I-FineEval: Fine-Grained Compositional Metric for Text-to-Image Evaluation

Seyed Mohammad Hadi Hosseini, Amir Mohammad Izadi, Ali Abdollahi +2

Although recent text-to-image generative models have achieved impressive performance, they still often struggle with capturing the compositional complexities of prompts including a…

cs.CV2025

Fine-Grained Alignment and Noise Refinement for Compositional Text-to-Image Generation

Amir Mohammad Izadi, Seyed Mohammad Hadi Hosseini, Soroush Vafaie Tabar +3

Text-to-image generative models have made significant advancements in recent years; however, accurately capturing intricate details in textual prompts-such as entity missing, attri…