1 paper · 1 filter
Romain Lacombe, Kerrie Wu, Eddie Dilworth
Large Language Models deployed as question answering tools require robust calibration to avoid overconfidence. We systematically evaluate how reasoning capabilities and budget affe…