3 papers
cs.AI2026
MirrorCode: AI can rebuild entire programs from behavior alone
Tom Adamczewski, David Owen, David Rein +4
AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C compiler. However, existing coding bench…
cs.CL2025
Large Language Models Often Know When They Are Being Evaluated
Joe Needham, Giles Edkins, Govind Pimpale +2
If AI models can detect when they are being evaluated, the effectiveness of evaluations might be compromised. For example, models could have systematically different behavior durin…
cs.LG2025
Evaluating Defences against Unsafe Feedback in RLHF
Domenic Rosati, Giles Edkins, Harsh Raj +5
While there has been progress towards aligning Large Language Models (LLMs) with human values and ensuring safe behaviour at inference time, safety guards can easily be removed whe…