10 papers
SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills
Nizhang Li, Zonghao Ying, Xiangfan Wu +7
External skills extend the capabilities of large language model agents, but also introduce an execution-time attack surface: a skill that appears benign under inspection may reveal…
SafeFlow: Semantic Information-Flow Control for Blocking Malicious Propagation in Multi-Agent Systems
Haowen Dai, Zonghao Ying, Wenfeng Li +10
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective c…
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
Dongdong Yang, Deyue Zhang, Zhao Liu +5
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulnerable to adversarial misuse, particularly…
ElephantAgent: Contextual State Continuity in Agentic Systems
Jiankai Jin, Xiangzheng Zhang, Zhao Liu +4
Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependencies introduce novel attack surfaces. Recen…
Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions
Shenao Wang, Xinyi Hou, Zhao Liu +5
GitHub Actions is increasingly used to deploy LLM-based agents for repository-centric tasks such as issue triage, pull-request review, code modification, and release assistance. Th…
SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
Zhe Liu, Zonghao Ying, Wenxin Zhang +5
Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and tool execution. While these capabilit…