11 papers
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Yechao Zhang, Shiqian Zhao, Jiawen Zhang +5
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. Thi…
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
Xianglin Yang, Bryan Hooi, Gelei Deng +2
The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulnerability. In this work, we intr…
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
Xuanye Zhang, Yongsen Zheng, Zhuqin Xu +5
LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents toward inappropriate/wrong too…
Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
Yechao Zhang, Shiqian Zhao, Jie Zhang +5
We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute…
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
Xianglin Yang, Gelei Deng, Jieming Shi +2
Large language models (LLMs) are vital for a wide range of applications yet remain susceptible to jailbreak threats, which could lead to the generation of inappropriate responses.…
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
Ruozhao Yang, Mingfei Cheng, Gelei Deng +3
Penetration testing is essential for assessing and strengthening system security against real-world threats, yet traditional workflows remain highly manual, expertise-intensive, an…