2 papers
cs.CL2026
ExCAM: Explainable Cultural Awareness Metrics
Christoph Leiter, Haiyue Song, Hour Kaing +4
Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent ben…
cs.CL2026
XQ-MEval: A Dataset with Cross-lingual Parallel Quality for Benchmarking Translation Metrics
Jingxuan Liu, Zhi Qu, Jin Tei +3
Automatic evaluation metrics are essential for building multilingual translation systems. The common practice of evaluating these systems is averaging metric scores across language…