Showing cs.CRShow all
3 papers · 1 filter
cs.CR2026
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
Chang Jin, An Wang, Zeming Wei +7
Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environm…
cs.CR2026
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
Chengcan Wu, Zhixin Zhang, Mingqian Xu +2
Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent design in MAS introduces unique trustworthi…
cs.CR2026
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction
Zeming Wei, Chengcan Wu, Meng Sun
Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In particular, they pose risks of generat…