binary reverse engineering 1dynamic debugging 1large language model agents 1vulnerability analysis 1zero-day discovery 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CR2026
Agentic Vulnerability Reasoning on COTS Binaries
Hwiwon Lee, Jongseong Kim, Lingming Zhang
The paper introduces SLYP, a REACT-style pipeline that enables large language model agents to discover and validate vulnerabilities directly in commercial off‑the‑shelf (COTS) Wind…
cs.CR2026
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
Hwiwon Lee, Jiawei Liu, Dongjun Kim +4
Finding a real vulnerability in complicated systems is a challenging, long-horizon task that demands reasoning across an entire codebase to produce a working proof-of-concept (PoC)…
cs.LG2025
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
Hwiwon Lee, Ziqi Zhang, Hanxiao Lu +1
Rigorous security-focused evaluation of large language model (LLM) agents is imperative for establishing trust in their safe deployment throughout the software development lifecycl…