4 papers · 1 filter
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
Yuhang Wang
Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit. We identify the plan…
Null Space Constrained Contrastive Visual Forgetting for MLLM Unlearning
Yuhang Wang, Zhenxing Niu, Haoxuan Ji +3
The core challenge of machine unlearning is to strike a balance between target knowledge removal and non-target knowledge retention. In the context of Multimodal Large Language Mod…
ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models
Yuhang Wang, Wenjie Mei, Junkai Zhang +3
Privacy deletion requests often arrive sequentially, creating a continual unlearning challenge for deployed multimodal large language models (MLLMs). However, existing benchmarks m…
From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
Yuhang Wang, Feiming Xu, Zheng Lin +6
Although large language model (LLM)-based agents, exemplified by OpenClaw, are increasingly evolving from task-oriented systems into personalized AI assistants for solving complex…