1 paper
Zihao Yang, Zijia Wang, Zhiqiu Huang
Multimodal large language models often capture visual-linguistic correlations but struggle to predict how local visual interventions propagate and affect downstream answers. We int…