1 paper · 1 filter
Johnathan Xie, Annie S. Chen, Yoonho Lee +2
The effectiveness of large language models (LLMs) is not only measured by their ability to generate accurate outputs but also by their calibration-how well their confidence scores…