1 citations · 1 across the 3 of their papers we have counts for
4 papers
GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events
Reworr, Artem Petrov, Dmitrii Volkov
OpenAI and DeepMind's AIs recently got gold at the IMO math olympiad and ICPC programming competition. We show frontier AI is similarly good at hacking by letting GPT-5 compete in…
Evaluating AI cyber capabilities with crowdsourced elicitation
Artem Petrov, Dmitrii Volkov
As AI systems become increasingly capable, understanding their offensive cyber potential is critical for informed governance and responsible deployment. However, it's hard to accur…
Demonstrating specification gaming in reasoning models
Alexander Bondarenko, Denis Volk, Dmitrii Volkov +1
We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the bench…
Hacking CTFs with Plain Agents
Rustem Turtayev, Artem Petrov, Dmitrii Volkov +1
We saturate a high-school-level hacking benchmark with plain LLM agent design. Concretely, we obtain 95% performance on InterCode-CTF, a popular offensive security benchmark, using…