From the 1 of 8 linked papers with an AI index.
8 papers
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
Ej Zhou, Lucas Resck, Zheng Hui +1
The paper shows that large language model evaluators give systematically different scores to the same content in different languages, favoring lower‑resource languages, even though…
Personality Without Persons? A Psychometric Critique of Big Five Testing in Large Language Models
Kim Zierahn, Cristina Cachero, Anna Korhonen +1
Human personality inventories are increasingly used to characterize large language models (LLMs), compare systems, and inform downstream governance claims. Yet, these inventories w…
Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization
Beiduo Chen, Pingjun Hong, Ziyun Zhang +3
Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisions. We study whether large…
Building Community-Centred NLP Resources for Puno Quechua
Elwin Huaman, Adrian Gamarra Lafuente, Johanna Cordova +1
The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (…
Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
Lucas Resck, Isabelle Augenstein, Anna Korhonen
Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptatio…
LLMs Aren't Human: A Critical Perspective on LLM Personality
Kim Zierahn, Cristina Cachero, Anna Korhonen +1
A growing body of research examines personality traits in Large Language Models (LLMs), particularly in human-agent collaboration. Prior work has frequently applied the Big Five in…