1 paper
Liran Tal, Johannes Kloos, Arsenii Rudich +2
We ran 300 repeated vulnerability-finding scans to measure how repeatable agentic large language model (LLM) security review is on the same JavaScript code, prompt, and benchmark h…