1 paper · 1 filter
Muskan Saraf, Sajjad Rezvani Boroujeni, Justin Beaudry +2
Large language models (LLMs) are increasingly deployed as evaluators of text quality, yet the validity of their judgments remains underexplored. This study investigates systematic…