Showing physics.ed-phShow all
2 papers · 1 filter
physics.ed-ph2026
LLM-as-a-judge validity in physics assessment depends more on the task than the model
Will Yeadon, Tom Hardy, Paul Mackay +1
As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge…
physics.ed-ph2024★ 3 cited
Evaluating AI and Human Authorship Quality in Academic Writing through Physics Essays
Will Yeadon, Elise Agra, Oto-obong Inyang +2
This study evaluates short-form physics essay submissions, equally divided between student work submitted before the introduction of ChatGPT and those generated by OpenAI…