1 paper · 1 filter
Wonjun Jeong, Dongseok Kim, Taegkeun Whangbo
Large Language Models (LLMs) can achieve inflated scores on multiple-choice tasks by exploiting inherent biases in option positions or labels, rather than demonstrating genuine und…