10 papers
CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking
Nanda Rani, Kimberly Milner, Minghao Shao +9
Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objective. They measure exploitation effect…
Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark
Minghao Shao, Nanda Rani, Kimberly Milner +9
Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate t…
AI In Cybersecurity Education -- Scalable Agentic CTF Design Principles and Educational Outcomes
Haoran Xi, Minghao Shao, Kimberly Milner +11
Large language models are rapidly changing how learners acquire and demonstrate cybersecurity skills. However, when human--AI collaboration is allowed, educators still lack validat…
Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries
Md Raz, Venkata Sai Charan Putrevu, Meet Udeshi +3
AI-powered malware increasingly exploits cloud-hosted generative-AI services and large language models (LLMs) as analysis engines for reconnaissance and code generation. Simultaneo…
SHIELD: A Host-Independent Framework for Ransomware Detection using Deep Filesystem Features
Md Raz, Venkata Sai Charan Putrevu, Prashanth Krishnamurthy +2
Ransomware's escalating sophistication necessitates tamper-resistant, off-host detection solutions that capture deep disk activity beyond the reach of a compromised operating syste…
Binary Diff Summarization using Large Language Models
Meet Udeshi, Venkata Sai Charan Putrevu, Prashanth Krishnamurthy +4
Security of software supply chains is necessary to ensure that software updates do not contain maliciously injected code or introduce vulnerabilities that may compromise the integr…