most citedCybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents

1 citations · 1 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CR2026

Dynamic Cyber Ranges

Víctor Mayoral-Vilches, María Sanz-Gómez, Francesco Balassone +6

As LLM-driven agents advance in cybersecurity, Jeopardy CTF benchmarks are approaching saturation and cyber ranges, the natural next evaluation frontier, offer diminishing resistan…

cs.CR2026

Towards Cybersecurity Superintelligence: from AI-guided humans to human-guided AI

Víctor Mayoral-Vilches, Stefan Rass, Martin Pinzger +14

Cybersecurity superintelligence -- artificial intelligence exceeding the best human capability in both speed and strategic reasoning -- represents the next frontier in security. Th…

cs.CR2026

Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense

Víctor Mayoral-Vilches, María Sanz-Gómez, Francesco Balassone +6

AI-driven penetration testing now executes thousands of actions per hour but still lacks the strategic intuition humans apply in competitive security. To build cybersecurity superi…

cs.CR2025

Cybersecurity AI: The World's Top AI Agent for Security Capture-the-Flag (CTF)

Víctor Mayoral-Vilches, Luis Javier Navarrete-Lozano, Francesco Balassone +4

Are Capture-the-Flag competitions obsolete? In 2025, Cybersecurity AI (CAI) systematically conquered some of the world's most prestigious hacking competitions, achieving Rank #1 at…

cs.CR2025

Cybersecurity AI in OT: Insights from an AI Top-10 Ranker in the Dragos OT CTF 2025

Víctor Mayoral-Vilches, Luis Javier Navarrete-Lozano, Francesco Balassone +3

Operational Technology (OT) cybersecurity increasingly relies on rapid response across malware analysis, network forensics, and reverse engineering disciplines. We examine the perf…

cs.CR2025★ 1 cited

Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents

María Sanz-Gómez, Víctor Mayoral-Vilches, Francesco Balassone +3

Cybersecurity spans multiple interconnected domains, complicating the development of meaningful, labor-relevant benchmarks. Existing benchmarks assess isolated skills rather than i…