1 paper · 1 filter
Eitan Wagner, Yuli Slavutsky, Omri Abend
Although language model scores are often treated as probabilities, their reliability as probability estimators has mainly been studied through calibration, overlooking other aspect…