2 papers
cs.CL2025
The QCET Taxonomy of Standard Quality Criterion Names and Definitions for the Evaluation of NLP Systems
Anya Belz, Simon Mille, Craig Thomson
Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quali…
cs.HC2024
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
Anya Belz, Craig Thomson
This paper presents version 3.0 of the Human Evaluation Datasheet (HEDS). This update is the result of our experience using HEDS in the context of numerous recent human evaluation…