Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent
Weidi Luo, Qiming Zhang, Yihao Quan +5
Coding agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn in…
cs.AI2026
FORTIS: Benchmarking Over-Privilege in Agent Skills
Shawn Li, Chenxiao Yu, Han Wang +8
Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This layer is widely treated as…
cs.AI2026
ShieldNet: Network-Level Guardrails against Emerging Supply-Chain Injections in Agentic Systems
Zhuowen Yuan, Zhaorun Chen, Zhen Xiang +5
Existing research on LLM agent security mainly focuses on prompt injection and unsafe input/output behaviors. However, as agents increasingly rely on third-party tools and MCP serv…