1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
Vaibhav Mavi, Shubh Jaroria, Weiqi Sun
Reliability and failure detection of large language models (LLMs) is critical for their deployment in high-stakes, multi-step reasoning tasks. Prior work explores confidence estima…
cs.CL2023★ 1 cited
Retrieval-Augmented Chain-of-Thought in Semi-structured Domains
Vaibhav Mavi, Abulhair Saparov, Chen Zhao
Applying existing question answering (QA) systems to specialized domains like law and finance presents challenges that necessitate domain expertise. Although large language models…