1 paper
Xunzhi Wang, Zhuowei Zhang, Gaonan Chen +7
Despite recent progress in systematic evaluation frameworks, benchmarking the uncertainty of large language models (LLMs) remains a highly challenging task. Existing methods for be…