2 papers
cs.CL2026
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Gaetano Perrone, Simon Pietro Romano
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and para…
cs.CR2025
Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs
Francesco Balassone, VÃctor Mayoral-Vilches, Stefan Rass +4
We empirically evaluate whether AI systems are more effective at attacking or defending in cybersecurity. Using CAI (Cybersecurity AI)'s parallel execution framework, we deployed a…