2 papers
cs.CL2026
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
Yutong Zhang, Jianshuo Dong, Peng Xu +5
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…
cs.CR2026
Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
Jianshuo Dong, Yiming Liu, Maosen Zhang +6
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this t…