2 citations · 2 across the 2 of their papers we have counts for
3 papers
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
Umesh Bodhwani, Thanh Tran, Kai Wei
Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-…
A Calibrated Reflection Approach for Enhancing Confidence Estimation in LLMs
Umesh Bodhwani, Yuan Ling, Shujing Dong +3
A critical challenge in deploying Large Language Models (LLMs) is developing reliable mechanisms to estimate their confidence, enabling systems to determine when to trust model out…
LentEx: Generalizable Latent Entity Extraction via Synthetic Data and Instruction-Tuned LLMs
Umesh Bodhwani, Yuan Ling, Cibi Chakravarthy Senthilkumar +4
Latent entity extraction (LEE) tackles the challenge of identifying implicit, contextually inferred entities within free text-an area where traditional entity extraction methods fa…