2 papers
cs.AI2025
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
Igor Ivanov
In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs…
cs.LG2025
Resurrecting saturated LLM benchmarks with adversarial encoding
Igor Ivanov, Dmitrii Volkov
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, o…