12 papers
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
Wanying Qu, Qinghua Mao, Yu Li +12
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime contro…
SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Siyuan Li, Aodu Wulianghai, Zehao Liu +8
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this p…
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks
Siyuan Li, Zehao Liu, Haoyu Li +5
As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reac…
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
Dongrui Liu, Yu Li, Zhonghao Yang +47
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI mod…
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
Jinze Gu, Qinghua Mao, Xi Lin +1
Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Graph RAG integrates knowledge gra…
Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
Qinghua Mao, Xi Lin, Jinze Gu +3
Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces critical safety risks: adversarie…