From the 1 of 13 linked papers with an AI index.
13 papers
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Jiawei Liang, Jianjie Huang, Ruoyu Chen +4
The paper introduces ERCR, a framework that improves token-level visual attribution in multimodal large language models by recombining evidence across multiple token-to-region view…
R-PGA: Robust Physical Adversarial Camouflage Generation via Relightable 3D Gaussian Splatting
Tianrui Lou, Siyuan Liang, Jiawei Liang +2
Physical adversarial camouflage poses a severe security threat to autonomous driving systems by mapping adversarial textures onto 3D objects. Nevertheless, current methods remain b…
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
Ruoyu Chen, Xiaoqing Guo, Kangwei Liu +6
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated token…
Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives
Kuanrong Liu, Siyuan Liang, Cheng Qian +2
As a general-purpose vision-language pretraining model, CLIP demonstrates strong generalization ability in image-text alignment tasks and has been widely adopted in downstream appl…
Text Adversarial Attacks with Dynamic Outputs
Wenqiang Wang, Siyuan Liang, Xiao Yan +1
Text adversarial attack methods are typically designed for static scenarios with fixed numbers of output labels and a predefined label space, relying on extensive querying of the v…
Poison Once, Control Anywhere: Clean-Text Visual Backdoors in VLM-based Mobile Agents
Xuan Wang, Siyuan Liang, Zhe Liu +5
Mobile agents powered by vision-language models (VLMs) are increasingly adopted for tasks such as UI automation and camera-based assistance. These agents are typically fine-tuned u…