2 papers
cs.CR2026
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
Hongwei Yao, Yiming Liu, Meihui Chen +6
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavioral safety and introduce ActBe…
cs.AI2026
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
Blaise Delattre, Cong Wang, Yang Cao
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the observed…