collaborators

5 papers

cs.CL2026

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

Yutong Zhang, Jianshuo Dong, Peng Xu +5

As large language models (LLMs) are deployed as autonomous agents, safety failures increasingly involve consequential actions. We study agentic misalignment, where agents take harm…

cs.CR2026

Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States

Jianshuo Dong, Yiming Liu, Maosen Zhang +6

Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this t…

cs.CR2026

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

Maosen Zhang, Jianshuo Dong, Boting Lu +5

LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. However, processing these contexts alongside…

cs.CR2025

DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling

Boheng Li, Junjie Wang, Yiming Li +7

Despite the integration of safety alignment and external filters, text-to-image (T2I) generative systems are still susceptible to producing harmful content, such as sexual or viole…

cs.CR2025

Can Large Language Models Automate the Refinement of Cellular Network Specifications?

Jianshuo Dong, Yuanjie Li, Jun Liu +2

Cellular networks, e.g., 4G/5G, rely on complex technical specifications to ensure correct functionality; however, these specifications often contain flaws or ambiguities. In this…