activity
20242026
collaborators

6 papers

cs.CR2026

Language Models Can Autonomously Hack and Self-Replicate

Alena Air, Reworr, Nikolaj Kotov +3

We demonstrate that language models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts. The agent independently finds and exploits…

cs.CR2025

GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events

Reworr, Artem Petrov, Dmitrii Volkov

OpenAI and DeepMind's AIs recently got gold at the IMO math olympiad and ICPC programming competition. We show frontier AI is similarly good at hacking by letting GPT-5 compete in…

cs.AI2025

Demonstrating specification gaming in reasoning models

Alexander Bondarenko, Denis Volk, Dmitrii Volkov +1

We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the bench…

cs.CR2025

Evaluating AI cyber capabilities with crowdsourced elicitation

Artem Petrov, Dmitrii Volkov

As AI systems become increasingly capable, understanding their offensive cyber potential is critical for informed governance and responsible deployment. However, it's hard to accur…

cs.CR2025

LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild

Reworr, Dmitrii Volkov

Attacks powered by Large Language Model (LLM) agents represent a growing threat to modern cybersecurity. To address this concern, we present LLM Honeypot, a system designed to moni…

cs.CR2024

Hacking CTFs with Plain Agents

Rustem Turtayev, Artem Petrov, Dmitrii Volkov +1

We saturate a high-school-level hacking benchmark with plain LLM agent design. Concretely, we obtain 95% performance on InterCode-CTF, a popular offensive security benchmark, using…