1 paper · 1 filter
Silvia Cappelletti, Tobia Poppi, Samuele Poppi +5
Large Language Models (LLMs) are increasingly evaluated on multiple-choice question answering (MCQA) tasks using *first-token probability* (FTP), which selects the answer option wh…