1 paper
Vaibhav Mavi, Shubh Jaroria, Weiqi Sun
Reliability and failure detection of large language models (LLMs) is critical for their deployment in high-stakes, multi-step reasoning tasks. Prior work explores confidence estima…