4 papers
Transferable Self-Evolving Playbooks for Agentic Security Auditing
Ziyue Wang, Cheuk Wang Maurice Ng, Chenchen Yu +3
An LLM agent for vulnerability discovery and validation is more than a model. It combines three components: an LLM for code analysis, an agent harness such as Codex or OpenCode for…
TxRay: Agentic Postmortem of Live Blockchain Attacks
Ziyue Wang, Jiangshan Yu, Kaihua Qin +3
Decentralized Finance (DeFi) has turned blockchains into financial infrastructure, allowing anyone to trade, lend, and build protocols without intermediaries, but this openness exp…
Agentic Discovery and Validation of Android App Vulnerabilities
Ziyue Wang, Liyi Zhou
Existing Android vulnerability detection tools overwhelm teams with thousands of low-signal warnings yet uncover few true positives. Analysts spend days triaging these results, cre…
Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities
Andrey Anurin, Jonathan Ng, Kibo Schaffer +2
LLM agents have the potential to revolutionize defensive cyber operations, but their offensive capabilities are not yet fully understood. To prepare for emerging threats, model dev…