13 papers
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Yechao Zhang, Shiqian Zhao, Jiawen Zhang +5
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. Thi…
Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
Yechao Zhang, Shiqian Zhao, Jie Zhang +5
We identify a critical security vulnerability in mainstream Claw personal AI agents: untrusted content encountered during heartbeat-driven background execution can silently pollute…
Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models
Jinman Wu, Yi Xie, Shen Lin +2
Safety alignment is often conceptualized as a monolithic process wherein harmfulness detection automatically triggers refusal. However, the persistence of jailbreak attacks suggest…
Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads
Jinman Wu, Yi Xie, Shiqian Zhao +1
Currently, open-sourced large language models (OSLLMs) have demonstrated remarkable generative performance. However, as their structure and weights are made public, they are expose…
When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems
Shiqian Zhao, Jiayang Liu, Yiming Li +9
Modern text-to-image (T2I) generation systems (e.g., DALLE 3) exploit the memory mechanism, which captures key information in multi-turn interactions for faithful generation…
Inference-time Alignment via Sparse Junction Steering
Runyi Hu, Jie Zhang, Shiqian Zhao +7
Token-level steering has emerged as a pivotal approach for inference-time alignment, enabling fine grained control over large language models by modulating their output distributio…