1 paper
Veronika Batzdorfer, Carlo Romano Marcello Alessandro Santagiustina
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset f…