6 papers
Language Models Can Autonomously Hack and Self-Replicate
Alena Air, Reworr, Nikolaj Kotov +3
We demonstrate that language models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts. The agent independently finds and exploits…
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
Reworr, Artem Petrov, Dmitrii Volkov
OpenAI and DeepMind's AIs recently got gold at the IMO math olympiad and ICPC programming competition. We show frontier AI is similarly good at hacking by letting GPT-5 compete in…
Demonstrating specification gaming in reasoning models
Alexander Bondarenko, Denis Volk, Dmitrii Volkov +1
We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the bench…
Evaluating AI cyber capabilities with crowdsourced elicitation
Artem Petrov, Dmitrii Volkov
As AI systems become increasingly capable, understanding their offensive cyber potential is critical for informed governance and responsible deployment. However, it's hard to accur…
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
Reworr, Dmitrii Volkov
Attacks powered by Large Language Model (LLM) agents represent a growing threat to modern cybersecurity. To address this concern, we present LLM Honeypot, a system designed to moni…
Hacking CTFs with Plain Agents
Rustem Turtayev, Artem Petrov, Dmitrii Volkov +1
We saturate a high-school-level hacking benchmark with plain LLM agent design. Concretely, we obtain 95% performance on InterCode-CTF, a popular offensive security benchmark, using…