activity
20232026
most citedGetting pwn'd by AI: Penetration Testing with Large Language Models

126 citations · 174 across the 16 of their papers we have counts for

collaborators
Showing cs.CRShow all

8 papers · 1 filter

cs.CR2026

The Ethics of Autonomous AI Agents for Offensive Security

Andreas Happe, Jürgen Cito, Jasmin Wachter

LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling - deterministic, narrowly scoped, and operated by trained practitioner…

cs.CR2026

Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents

Benjamin Probst, Andreas Happe, Jürgen Cito

Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise security, privacy, and sovereignty conce…

cs.CR2025★ 1 cited

On the Surprising Efficacy of LLMs for Penetration-Testing

Andreas Happe, Jürgen Cito

This paper presents a critical examination of the surprising efficacy of Large Language Models (LLMs) in penetration testing. The paper thoroughly reviews the evolution of LLMs and…

cs.CR2025

Recognition Without Mitigation: Ethical Frameworks in Autonomous Offensive-LLM Agent Research

Andreas Happe, Jürgen Cito

Large language models have moved from advising on offensive security to autonomously conducting it. A growing literature presents agents that execute reconnaissance, exploitation,…

cs.CR2025★ 2 cited

Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design

Andreas Happe, Jürgen Cito

Large Language Models (LLMs) have emerged as a powerful approach for driving offensive penetration-testing tooling. Due to the opaque nature of LLMs, empirical methods are typicall…

cs.CR2025★ 10 cited

Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks

Andreas Happe, Jürgen Cito

Enterprise penetration-testing is often limited by high operational costs and the scarcity of human expertise. This paper investigates the feasibility and effectiveness of using La…