10 papers
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
Wanying Qu, Qinghua Mao, Yu Li +12
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime contro…
SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
Siyuan Li, Aodu Wulianghai, Zehao Liu +8
Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this p…
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
Dongrui Liu, Yu Li, Zhonghao Yang +47
Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI mod…
GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction
Jinze Gu, Qinghua Mao, Xi Lin +1
Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Graph RAG integrates knowledge gra…
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
Leitao Yuan, Qinghua Mao, Daizong Liu +5
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate encoders can generalize to clo…
Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
Qinghua Mao, Xi Lin, Jinze Gu +3
Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces critical safety risks: adversarie…