2 papers
cs.CL2024
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
Xinpeng Wang, Bolei Ma, Chengzhi Hu +5
The open-ended nature of language generation makes the evaluation of autoregressive large language models (LLMs) challenging. One common evaluation approach uses multiple-choice qu…
cs.CL2024
VariErr NLI: Separating Annotation Error from Human Label Variation
Leon Weber-Genzel, Siyao Peng, Marie-Catherine de Marneffe +1
Human label variation arises when annotators assign different labels to the same item for valid reasons, while annotation errors occur when labels are assigned for invalid reasons.…