2 papers
cs.AI2026
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
Litian Zhang, Chaozhuo Li, Yuting Zhang +3
LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions through inter-agent communication…
cs.CR2026
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
Songyang Liu, Chaozhuo Li, Rui Pu +5
Jailbreak attacks present a significant challenge to the safety of Large Language Models (LLMs), yet current automated evaluation methods largely rely on coarse classifications tha…