3 papers
cs.CL2026
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Gaetano Perrone, Simon Pietro Romano
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and para…
cs.CR2025
Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs
Francesco Balassone, VÃctor Mayoral-Vilches, Stefan Rass +4
We empirically evaluate whether AI systems are more effective at attacking or defending in cybersecurity. Using CAI (Cybersecurity AI)'s parallel execution framework, we deployed a…
cs.CR2025
Hybrid Privilege Escalation and Remote Code Execution Exploit Chains
Miguel Tulla, Andrea Vignali, Christian Colon +5
Research on exploit chains predominantly focuses on sequences with one type of exploit, e.g., either escalating privileges on a machine or executing remote code. In networks, hybri…