4 papers
INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment
Yutong Zhang, Jianshuo Dong, Peng Xu +5
As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…
Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States
Jianshuo Dong, Yiming Liu, Maosen Zhang +6
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this t…
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
Maosen Zhang, Jianshuo Dong, Boting Lu +5
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside…
AlphaResearch: Accelerating New Algorithm Discovery with Language Models
Zhaojian Yu, Kaiyue Feng, Yilun Zhao +3
LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch},…