2 papers
cs.AI2025
A Rosetta Stone for AI Benchmarks
Anson Ho, Jean-Stanislas Denain, David Atanasov +2
Most AI benchmarks saturate within years or even months after they are introduced, making it hard to study long-run trends in AI capabilities. To address this challenge, we build a…
cs.LG2025
Evaluating Defences against Unsafe Feedback in RLHF
Domenic Rosati, Giles Edkins, Harsh Raj +5
While there has been progress towards aligning Large Language Models (LLMs) with human values and ensuring safe behaviour at inference time, safety guards can easily be removed whe…