2 papers
cs.AI2026
The Answer Is Not the Argument
Will Yeadon, Sergio Juárez, Paul Mackay +5
Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves reasoning ve…
physics.ed-ph2026
LLM-as-a-judge validity in physics assessment depends more on the task than the model
Will Yeadon, Tom Hardy, Paul Mackay +1
As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge…