works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.CV2026

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

Jiawei Liang, Jianjie Huang, Ruoyu Chen +4

The paper introduces ERCR, a framework that improves token-level visual attribution in multimodal large language models by recombining evidence across multiple token-to-region view…

cs.CV2026

R-PGA: Robust Physical Adversarial Camouflage Generation via Relightable 3D Gaussian Splatting

Tianrui Lou, Siyuan Liang, Jiawei Liang +2

Physical adversarial camouflage poses a severe security threat to autonomous driving systems by mapping adversarial textures onto 3D objects. Nevertheless, current methods remain b…

cs.CV2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…

cs.CV2025

Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives

Kuanrong Liu, Siyuan Liang, Cheng Qian +2

As a general-purpose vision-language pretraining model, CLIP demonstrates strong generalization ability in image-text alignment tasks and has been widely adopted in downstream appl…

cs.CV2025

Text Adversarial Attacks with Dynamic Outputs

Wenqiang Wang, Siyuan Liang, Xiao Yan +1

Text adversarial attack methods are typically designed for static scenarios with fixed numbers of output labels and a predefined label space, relying on extensive querying of the v…

cs.CR2025

Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents

Xuan Wang, Siyuan Liang, Zhe Liu +5

Mobile agents powered by vision-language models (VLMs) are increasingly adopted for tasks such as UI automation and camera-based assistance. These agents are typically fine-tuned u…