works on

From the 2 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

Xianghao Jiao, Ruoyu Chen, Wei Wang +6

Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inco…

cs.RO2026

Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving

Yun Li, Jiachen Gong, Simon Thompson +7

Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the control rate a vehicle requires. Existing…

cs.CV2026

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

Jiawei Liang, Jianjie Huang, Ruoyu Chen +4

The paper introduces ERCR, a framework that improves token-level visual attribution in multimodal large language models by recombining evidence across multiple token-to-region view…

cs.CV2026

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8

The paper introduces a training approach that encodes human‑provided region priors and uses a subset‑selection attribution method to penalize models when their decision evidence fa…

cs.CV2026

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

Yannan Chen, Ruoyu Chen, Wei Wang +6

Current visual models often make predictions based on a limited set of discriminative visual cues. As a result, they may become unreliable when the distribution shifts or when thes…

cs.CV2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…