2 papers
cs.CR2026
Transferable Self-Evolving Playbooks for Agentic Security Auditing
Ziyue Wang, Cheuk Wang Maurice Ng, Chenchen Yu +3
An LLM agent for vulnerability discovery and validation is more than a model. It combines three components: an LLM for code analysis, an agent harness such as Codex or OpenCode for…
cs.AI2026
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
Strick Sheng, Ziyue Wang, Liyi Zhou
Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tra…