3 papers
cs.CR2026
Language Models Can Autonomously Hack and Self-Replicate
Alena Air, Reworr, Nikolaj Kotov +3
We demonstrate that language models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts. The agent independently finds and exploits…
cs.CR2025
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
Reworr, Artem Petrov, Dmitrii Volkov
OpenAI and DeepMind's AIs recently got gold at the IMO math olympiad and ICPC programming competition. We show frontier AI is similarly good at hacking by letting GPT-5 compete in…
cs.CR2025
LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild
Reworr, Dmitrii Volkov
Attacks powered by Large Language Model (LLM) agents represent a growing threat to modern cybersecurity. To address this concern, we present LLM Honeypot, a system designed to moni…