3 papers
cs.CR2026
Lessons from Penetration Tests on Large-Scale Agent Systems
Kevin Eykholt, Dhilung Kirat, Xiaokui Shu +3
As AI systems gain increasing autonomy and execution capability, the number of discovered security vulnerabilities continues to rise. However, many of these vulnerabilities are not…
cs.CR2026
Understanding Human-AI Collaboration in Cybersecurity Competitions
Tingxuan Tang, Nicolas Janis, Kalyn Asher Montague +6
Capture-the-Flag (CTF) competitions are increasingly becoming a testbed for evaluating AI capabilities at solving security tasks, due to the controlled environments and objective s…
cs.SE2025
One Bug, Hundreds Behind: LLMs for Large-Scale Bug Discovery
Qiushi Wu, Yue Xiao, Dhilung Kirat +3
Fixing bugs in large programs is a challenging task that demands substantial time and effort. Once a bug is found, it is reported to the project maintainers, who work with the repo…