52 citations · 106 across the 28 of their papers we have counts for
3 papers · 1 filter
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
Kaustubh D. Dhole, Charles L. A. Clarke, Eugene Y. Agichtein
Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, requir…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
An Interactive Query Generation Assistant using LLM-based Prompt Modification and User Feedback
Kaustubh D. Dhole, Ramraj Chandradevan, Eugene Agichtein
While search is the predominant method of accessing information, formulating effective queries remains a challenging task, especially for situations where the users are not familia…