1 paper
Muskan Saraf, Sajjad Rezvani Boroujeni, Justin Beaudry +2
Large language models (LLMs) are increasingly deployed as evaluators of text quality, yet the validity of their judgments remains underexplored. This study investigates systematic…