activity
20242026
collaborators

12 papers

cs.CR2026

ReCon: A Resource-Constrained Benchmark for LLM-Based Cybersecurity Compliance Across Ingestion and Retrieval Pipelines

Rohit Negi, Rishik Jain, Soumyo V Chakarborty +2

With the increasingly aggressive cyber threat landscape for governments, businesses, and institutions, as information and/or cybersecurity implementations are increasingly under sc…

cs.CR2026

CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking

Nanda Rani, Kimberly Milner, Minghao Shao +9

Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effect…

cs.CR2026

An Automated Framework for Cybersecurity Policy Compliance Assessment Against Security Control Standards

Bikash Saha, Sandeep Kumar Shukla

Organizational cybersecurity policies are often examined to determine whether they adequately comply standard security controls. This task is difficult because control statements a…

cs.CR2026

MalGEN: A Testbed for Modeling and Evaluating Malware Behaviors

Bikash Saha, Sandeep Kumar Shukla

Modern cybersecurity requires systematic ways to evaluate how detection systems respond to evolving and previously unseen attack behaviors. Existing malware repositories largely ca…

cs.CR2026

Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark

Minghao Shao, Nanda Rani, Kimberly Milner +9

Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate t…

cs.SE2026

AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes

Haoran Xi, Minghao Shao, Kimberly Milner +11

Large language models are rapidly changing how learners acquire and demonstrate cybersecurity skills. However, when human--AI collaboration is allowed, educators still lack validat…