12 papers
ReCon: A Resource-Constrained Benchmark for LLM-Based Cybersecurity Compliance Across Ingestion and Retrieval Pipelines
Rohit Negi, Rishik Jain, Soumyo V Chakarborty +2
With the increasingly aggressive cyber threat landscape for governments, businesses, and institutions, as information and/or cybersecurity implementations are increasingly under sc…
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Nanda Rani, Kimberly Milner, Minghao Shao +9
Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effect…
An Automated Framework for Cybersecurity Policy Compliance Assessment Against Security Control Standards
Bikash Saha, Sandeep Kumar Shukla
Organizational cybersecurity policies are often examined to determine whether they adequately comply standard security controls. This task is difficult because control statements a…
MalGEN: A Testbed for Modeling and Evaluating Malware Behaviors
Bikash Saha, Sandeep Kumar Shukla
Modern cybersecurity requires systematic ways to evaluate how detection systems respond to evolving and previously unseen attack behaviors. Existing malware repositories largely ca…
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
Minghao Shao, Nanda Rani, Kimberly Milner +9
Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate t…
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
Haoran Xi, Minghao Shao, Kimberly Milner +11
Large language models are rapidly changing how learners acquire and demonstrate cybersecurity skills. However, when human--AI collaboration is allowed, educators still lack validat…