works on

From the 2 of 16 linked papers with an AI index.

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

Xianghao Jiao, Ruoyu Chen, Wei Wang +6

Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inco…

cs.CV2026

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

Jiawei Liang, Jianjie Huang, Ruoyu Chen +4

The paper introduces ERCR, a framework that improves token-level visual attribution in multimodal large language models by recombining evidence across multiple token-to-region view…

cs.CV2026

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8

The paper introduces a training approach that encodes human‑provided region priors and uses a subset‑selection attribution method to penalize models when their decision evidence fa…

cs.CV2026

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

Yannan Chen, Ruoyu Chen, Wei Wang +6

Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…

cs.CV2026

Domain Adaptive Object Detection via Dual-Stream Bilevel-Cycle Optimization

Yannan Chen, Wei Wang, Wenqiang Wang +5

Cycle self-training (CST) breaks the shared classifier assumption of the standard self-training framework, which is effective for unsupervised domain adaptation and exploits unlabe…

cs.CV2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…