5 papers
Building Comparative Motivation Profiles with Instrumental Interventions
David Vella Zarb, Rustem Turtayev, Taywon Min +2
Safety evaluations often infer latent motivations from behavioral patterns, but the construct validity of these inferences is unclear. We study this problem in alignment faking, wh…
CayleyPy RL: Pathfinding and Reinforcement Learning on Cayley Graphs
A. Chervov, M. Obozov, A. Soibelman +31
This paper is the second in a series of studies on developing efficient artificial intelligence-based approaches to pathfinding on extremely large graphs (e.g. nodes) wit…
CayleyPy Growth: Efficient growth computations and hundreds of new conjectures on Cayley graphs (Brief version)
A. Chervov, D. Fedoriaka, E. Konstantinova +46
This is the third paper of the CayleyPy project applying artificial intelligence to problems in group theory. We announce the first public release of CayleyPy, an open source Pytho…
Misalignment Bounty: Crowdsourcing AI Agent Misbehavior
Rustem Turtayev, Natalia Fedorova, Oleg Serikov +3
Advanced AI systems sometimes act in ways that differ from human intent. To gather clear, reproducible examples, we ran the Misalignment Bounty: a crowdsourced project that collect…
Hacking CTFs with Plain Agents
Rustem Turtayev, Artem Petrov, Dmitrii Volkov +1
We saturate a high-school-level hacking benchmark with plain LLM agent design. Concretely, we obtain 95% performance on InterCode-CTF, a popular offensive security benchmark, using…