3 papers
cs.CR2026
Twin Agent: Context Residual Compression for Privilege Separated Agents
Zhanhao Hu, Dennis Jacob, Xiao Huang +3
Large language model (LLM) agents are vulnerable to security risks, such as prompt injection attacks from untrusted context that manipulate downstream reasoning and tool use. Exist…
cs.CR2025
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Julien Piet, Xiao Huang, Dennis Jacob +7
Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulne…
cs.CR2025
PromptShield: Deployable Detection for Prompt Injection Attacks
Dennis Jacob, Hend Alzahrani, Zhanhao Hu +2
Application designers have moved to integrate large language models (LLMs) into their products. However, many LLM-integrated applications are vulnerable to prompt injections. While…