6 papers
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye +48
As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that st…
Medical thinking with multiple images
Zonghai Yao, Benlu Wang, Yifan Zhang +8
Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…
Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics
Jinu Lee, Kyoung-Woon On, Simeng Han +2
Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the…
REVERE: Reflective Evolving Research Engineer
Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan +1
Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on weak update mechanisms, such as full-prompt…
Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights
Hyunjae Kim, Jiwoong Sohn, Aidan Gilson +24
Large language models (LLMs) are transforming the landscape of medicine, yet two fundamental challenges persist: keeping up with rapidly evolving medical knowledge and providing ve…
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
Yicheng Gao, Gonghan Xu, Zhe Wang +1
Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evalua…