32 citations · 125 across the 39 of their papers we have counts for
1 paper · 2 filters
Alessandro Bondielli, Lucia Passaro, Davide Bacciu +1
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, st…