7 papers
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
Kaiyue Yang, Yuyan Bu, Jingwei Yi +5
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on saf…
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
Shi Liu, Xuehai Tang, Xikang Yang +4
The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented autonomous execution capabilit…
RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents
Wenjie Xiao, Xuehai Tang, Biyu Zhou +2
Agent skills introduce a new and more severe form of indirect injection for LLM agents: unlike traditional indirect prompt injection, attackers can hide malicious instructions insi…
FABLE: Fine-grained Fact Anchoring for Unstructured Model Editing
Peng Wang, Biyu Zhou, Xuehai Tang +2
Unstructured model editing aims to update models with real-world text, yet existing methods often memorize text holistically without reliable fine-grained fact access. To address t…
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
Xikang Yang, Biyu Zhou, Xuehai Tang +2
Large Language Models (LLMs) demonstrate impressive capabilities across a wide range of tasks, yet their safety mechanisms remain susceptible to adversarial attacks that exploit co…
LyapLock: Bounded Knowledge Preservation in Sequential Large Language Model Editing
Peng Wang, Biyu Zhou, Xuehai Tang +2
Large Language Models often contain factually incorrect or outdated knowledge, giving rise to model editing methods for precise knowledge updates. However, current mainstream locat…