1 paper · 2 filters
Eitan Wagner, Yuli Slavutsky, Omri Abend
Although language model scores are often treated as probabilities, their reliability as probability estimators has mainly been studied through calibration, overlooking other aspect…