From the 1 of 9 linked papers with an AI index.
9 papers
When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Yongli Xiang, Zhifang Zhang, Bojun Yang +4
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrat…
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
Zhifang Zhang, Qiqi Tao, Jiaqi Lv +3
The paper introduces TokenSwap, a stealthy backdoor attack on large vision-language models that swaps key textual tokens to corrupt the model's understanding of object relationship…
Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
Haochang Hao, Dehai Min, Zhifang Zhang +4
Agent skills extend general-purpose agents, but their open format enables skill poisoning: a tampered skill can make an agent run an attacker's command while completing the user's…
Towards Safer Large Reasoning Models by Promoting Safety Decision-Making before Chain-of-Thought Generation
Jianan Chen, Zhifang Zhang, Shuo He +3
Large reasoning models (LRMs) achieved remarkable performance via chain-of-thought (CoT), but recent studies showed that such enhanced reasoning capabilities are at the expense of…
Test-Time Attention Purification for Backdoored Large Vision Language Models
Zhifang Zhang, Bojun Yang, Shuo He +5
Despite the strong multimodal performance, large vision-language models (LVLMs) are vulnerable during fine-tuning to backdoor attacks, where adversaries insert trigger-embedded sam…
Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning
Zhifang Zhang, Shuo He, Haobo Wang +2
Multimodal contrastive learning models (e.g., CLIP) can learn high-quality representations from large-scale image-text datasets, while they exhibit significant vulnerabilities to b…