3 papers
cs.LG2025
JULI: Jailbreak Large Language Models by Self-Introspection
Jesson Wang, Zhanhao Hu, David Wagner
Large Language Models (LLMs) are trained with safety alignment to prevent generating malicious content. Although some attacks have highlighted vulnerabilities in these safety-align…
cs.CR2025
JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
Julien Piet, Xiao Huang, Dennis Jacob +7
Safety and security remain critical concerns in AI deployment. Despite safety training through reinforcement learning with human feedback (RLHF) [ 32], language models remain vulne…
cs.CR2025
PromptShield: Deployable Detection for Prompt Injection Attacks
Dennis Jacob, Hend Alzahrani, Zhanhao Hu +2
Application designers have moved to integrate large language models (LLMs) into their products. However, many LLM-integrated applications are vulnerable to prompt injections. While…