12 citations · 18 across the 3 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.CR2024★ 3 cited
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
Luca Gioacchini, Marco Mellia, Idilio Drago +3
Generative AI agents, software systems powered by Large Language Models (LLMs), are emerging as a promising approach to automate cybersecurity tasks. Among the others, penetration…
cs.AI2024★ 2 cited
AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
Luca Gioacchini, Giuseppe Siracusano, Davide Sanvito +4
The advances made by Large Language Models (LLMs) have led to the pursuit of LLM agents that can solve intricate, multi-step reasoning tasks. As with any research pursuit, benchmar…