2 papers
cs.AI2025
Demonstrating specification gaming in reasoning models
Alexander Bondarenko, Denis Volk, Dmitrii Volkov +1
We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the bench…
cs.CR2024
Hacking CTFs with Plain Agents
Rustem Turtayev, Artem Petrov, Dmitrii Volkov +1
We saturate a high-school-level hacking benchmark with plain LLM agent design. Concretely, we obtain 95% performance on InterCode-CTF, a popular offensive security benchmark, using…