collaborators

14 papers

cs.CR2026

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Lipeng He, Yihan Wang, Jiawen Zhang +1

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses…

cs.AI2026

Beyond Similarity: Trustworthy Memory Search for Personal AI Agents

Jiawen Zhang, Kejia Chen, Jiachen Ma +7

Personal AI agents increasingly rely on long-term memory to provide persistent personalization across sessions. However, existing memory pipelines are largely driven by semantic si…

cs.LG2026

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

Jiachen Ma, Jiawen Zhang, Xiangtian Li +3

While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circumvent conventional surface-…

cs.AI2026

Evaluating Cognitive Age Alignment in Interactive AI Agents

Yifan Shen, Jiawen Zhang, Jian Xu +4

While agentic AI and its core multimodal large language models (MLLMs) have demonstrated remarkable promise in language and visual reasoning across domains ranging from daily life…

cs.CR2026

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

Kejia Chen, Jiawen Zhang, Boheng Li +6

Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demonstrations. We study why this a…

cs.AI2026

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable

Kejia Chen, Jiawen Zhang, Yihong Wu +5

Large reasoning models often reach correct answers through flawed intermediate steps, creating a gap between final accuracy and reasoning reliability. Existing alignment strategies…