10 papers
Causal Attribution via Activation Patching
Amirmohammad Izadi, Mohammadali Banayeeanzade, Alireza Mirrokni +4
Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains ch…
Detecting Popular Social Events through Limited Observation with Deep Survival Analysis
Maryam Ramezani, Hossein Goli, AmirMohammad Izadi +1
Users increasing activity across various social networks made it the most widely used platform for exchanging and propagating information among individuals. To spread information w…
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
Hosein Hasani, Mohammadali Banayeeanzade, Ali Nafisi +5
Large language models (LLMs), despite strong performance on complex mathematical problems, exhibit systematic limitations in counting tasks. This issue arises from the architectura…
Understanding Counting Mechanisms in Large Language and Vision-Language Models
Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari +4
Counting is one of the fundamental abilities of large language models (LLMs) and large vision-language models (LVLMs). This paper examines how these foundation models represent and…
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari +4
Large vision-language models (LVLMs) show strong performance across multimodal benchmarks but remain limited in structured reasoning and precise grounding. Recent work has demonstr…
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
Mohammadjavad Ahmadpour, Amirmahdi Meighani, Payam Taebi +3
Test-time scaling (TTS) has emerged as a powerful paradigm for improving the reasoning ability of Large Language Models (LLMs) by allocating additional computation at inference, ye…