2 papers
cs.CL2026
The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents
Chen Chen, Kim Young Il, Yuan Yang +7
Large language model (LLM) agents with extended autonomy unlock new capabilities, but also introduce heightened challenges for LLM safety. In particular, an LLM agent may pursue ob…
cs.CR2025
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs
Xueluan Gong, Mingzhe Li, Yilin Zhang +5
Large Language Models (LLMs) have excelled in various tasks but are still vulnerable to jailbreaking attacks, where attackers create jailbreak prompts to mislead the model to produ…