From the 1 of 1 linked paper with an AI index.
1 paper
Ej Zhou, Lucas Resck, Zheng Hui +1
The paper shows that large language model evaluators give systematically different scores to the same content in different languages, favoring lower‑resource languages, even though…