2 papers
cs.CL2026
Individual and Combined Effects of English as a Second Language and Typos on LLM Performance
Serena Liu, Yutong Yang, Prisha Sheth +9
Large language models (LLMs) are used globally, and because much of their training data is in English, they typically perform best on English inputs. As a result, many non-native E…
cs.CL2026
Grading Scale Impact on LLM-as-a-Judge: Human-LLM Alignment Is Highest on 0-5 Grading Scale
Weiyue Li, Minda Zhao, Weixuan Dong +12
Large language models (LLMs) are increasingly used as automated evaluators, yet prior works demonstrate that these LLM judges often lack consistency in scoring when the prompt is a…