7 papers
Twin Agent: Context Residual Compression for Privilege Separated Agents
Zhanhao Hu, Dennis Jacob, Xiao Huang +3
Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Exist…
Test-time Adversarial Takeover: A Real-time Hijacking Interface against Robotic Diffusion Policies
Zi Yin, Peilin Chai, Siyuan Huang +1
Diffusion-based action generation has become a foundational component of embodied AI, but its reliance on visual conditioning leaves deployed visuomotor policies vulnerable to adve…
GradShield: Alignment Preserving Finetuning
Zhanhao Hu, Xiao Huang, Patrick Mendoza +4
Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implicitly harmful data. Even some…
Preventing Prompt Injection with Type-Directed Privilege Separation
Dennis Jacob, Emad Alghamdi, Zhanhao Hu +2
Modern language models have enabled the development of agentic systems that achieve strong performance on reasoning-intensive tasks. Unfortunately, this has come with a security co…
JULI: Jailbreak Large Language Models by Self-Introspection
Jesson Wang, Zhanhao Hu, David Wagner
Large Language Models (LLMs) are trained with safety alignment to prevent generating malicious content. Although some attacks have highlighted vulnerabilities in these safety-align…
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Julien Piet, Xiao Huang, Dennis Jacob +7
Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulne…