1 paper
Isaac Chung, Linda Freienthal
Cross-lingual evaluation of large language models (LLMs) typically conflates two sources of variance: genuine model performance differences and measurement instability. We investig…