1 paper
Anya Belz, Simon Mille, Craig Thomson
Prior work has shown that two NLP evaluation experiments that report results for the same quality criterion name (e.g. Fluency) do not necessarily evaluate the same aspect of quali…