2 papers
cs.CR2026
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
Alankrit Chona, Igor Kozlov, Ambuj Kumar
We introduce the Cyber Defense Benchmark, a benchmark for measuring how well large language model (LLM) agents perform the core SOC analyst task of threat hunting: given a database…
cs.AI2024
TaskGen: A Task-Based, Memory-Infused Agentic Framework using StrictJSON
John Chong Min Tan, Prince Saroj, Bharat Runwal +6
TaskGen is an open-sourced agentic framework which uses an Agent to solve an arbitrary task by breaking them down into subtasks. Each subtask is mapped to an Equipped Function or a…