3 papers
physics.ed-ph2026
LLM-as-a-judge validity in physics assessment depends more on the task than the model
Will Yeadon, Tom Hardy, Paul Mackay +1
As large language models (LLMs) are increasingly considered for automated assessment and feedback, understanding when LLM marking is valid is essential. We evaluate LLM-as-a-judge…
cs.LG2026
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
quant-ph2025
QSTToolkit: A Python Library for Deep Learning Powered Quantum State Tomography
George FitzGerald, Will Yeadon
We introduce QSTToolkit, a Python library for performing quantum state tomography (QST) on optical quantum state measurement data. The toolkit integrates traditional Maximum Likeli…