1 paper
Nisrine Rair, Alban Goupil, Valeriu Vrabie +1
Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity, especially when human annotators…