2 papers
cs.SE2026
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
Romain Gerard, Assmaa Zeghaider, Yan Guo
Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their capability remain poorly characte…
cs.CR2026
Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing
Liwei Yu, Shuo Li, Ming Zhou +2
Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptible to error cascading: failures…