10 papers
Transferable Self-Evolving Playbooks for Agentic Security Auditing
Ziyue Wang, Cheuk Wang Maurice Ng, Chenchen Yu +3
An LLM agent for vulnerability discovery and validation is more than a model. It combines three components: an LLM for code analysis, an agent harness such as Codex or OpenCode for…
Evasion Under Blockchain Sanctions
Endong Liu, Mark Ryan, Liyi Zhou +1
Sanctioning blockchain addresses has become a common regulatory response to malicious activities. However, enforcement on permissionless blockchains remains challenging due to comp…
When Agents Overtrust Environmental Evidence: An Extensible Agentic Framework for Benchmarking Evidence-Grounding Defects in LLM Agents
Strick Sheng, Ziyue Wang, Liyi Zhou
Large language model agents increasingly operate through environment-facing scaffolds that expose files, web pages, APIs, and logs. These observations influence tool use, state tra…
Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
Shanshan Gao, Liyi Zhou
Interactive agent benchmarks map an agent run to a binary outcome through outcome checks. When these checks rely on surface level signals or fail to capture the agent's actual acti…
TxRay: Agentic Postmortem of Live Blockchain Attacks
Ziyue Wang, Jiangshan Yu, Kaihua Qin +3
Decentralized Finance (DeFi) has turned blockchains into financial infrastructure, allowing anyone to trade, lend, and build protocols without intermediaries, but this openness exp…
AI Agent Smart Contract Exploit Generation
Arthur Gervais, Liyi Zhou
Smart contract vulnerabilities have led to billions in losses, yet finding actionable exploits remains challenging. Traditional fuzzers rely on rigid heuristics and struggle with c…