6 papers · 1 filter
Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations
Xianghao Jiao, Ruoyu Chen, Wei Wang +6
Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inco…
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints
Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8
Reliable models should not only predict correctly, but also base their decisions on acceptable evidence. However, conventional supervised learning typically provides only class-lev…
Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
Yannan Chen, Ruoyu Chen, Wei Wang +6
Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Jiawei Liang, Jianjie Huang, Xianghao Jiao +3
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens att…
Interpreting Object-level Foundation Models via Visual Precision Search
Ruoyu Chen, Siyuan Liang, Jingzhi Li +5
Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. Howev…