3 papers
cs.CR2026
Autonomous LLM Agents & CTFs: A Second Look
Youness Bouchari, Matteo Boffa, Marco Mellia +3
Large Language Model (LLM) agents are increasingly proposed to automate offensive security tasks, with recent studies reporting near human-level success rates in Capture-the-Flag (…
cs.CR2026
Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning
Jianan Huang, Rodolfo V. Valentim, Luca Vassio +4
The use of ML in cybersecurity has long been impaired by generalization issues: Models that work well in controlled scenarios fail to maintain performance in production. The root c…
cs.CR2026
CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics
Stefano Fumero, Kai Huang, Matteo Boffa +3
Post-mortem analysis of compromised systems is a key aspect of cyber forensics, today a mostly manual, slow, and error-prone task. Agentic AI, i.e., LLM-powered agents, is a promis…