activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

Xianghao Jiao, Ruoyu Chen, Wei Wang +6

Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inco…

cs.CV2026

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8

Reliable models should not only predict correctly, but also base their decisions on acceptable evidence. However, conventional supervised learning typically provides only class-lev…

cs.CV2025

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

Yannan Chen, Ruoyu Chen, Wei Wang +6

Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…

cs.CV2025

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…

cs.CV2025

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

Jiawei Liang, Jianjie Huang, Xianghao Jiao +3

Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens att…

cs.CV2024

Interpreting Object-level Foundation Models via Visual Precision Search

Ruoyu Chen, Siyuan Liang, Jingzhi Li +5

Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. Howev…