3 papers
cs.CL2026
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
Igor Ivanov, David Demitri Africa
Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignme…
cs.AI2025
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
Igor Ivanov
In this paper, LLMs are tasked with completing an impossible quiz, while they are in a sandbox, monitored, told about these measures and instructed not to cheat. Some frontier LLMs…
cs.LG2025
Resurrecting saturated LLM benchmarks with adversarial encoding
Igor Ivanov, Dmitrii Volkov
Recent work showed that small changes in benchmark questions can reduce LLMs' reasoning and recall. We explore two such changes: pairing questions and adding more answer options, o…