3 papers
cs.CL2026
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
Yutong Zhang, Jianshuo Dong, Peng Xu +5
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…
cs.CR2026
Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
Jianshuo Dong, Yiming Liu, Maosen Zhang +6
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this t…
cs.CR2023
One-bit Flip is All You Need: When Bit-flip Attack Meets Model Training
Jianshuo Dong, Han Qiu, Yiming Li +5
Deep neural networks (DNNs) are widely deployed on real-world devices. Concerns regarding their security have gained great attention from researchers. Recently, a new weight modifi…