activity
20242026
collaborators

6 papers

cs.LG2026

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Santiago Gonzalez, Alireza Amiri Bavandpour, Peter Ye +48

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that st…

cs.CV2026

Medical thinking with multiple images

Zonghai Yao, Benlu Wang, Yifan Zhang +8

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…

cs.AI2026

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

Jinu Lee, Kyoung-Woon On, Simeng Han +2

Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainability, yet remains challenging due to the…

cs.SE2026

REVERE: Reflective Evolving Research Engineer

Balaji Dinesh Gangireddi, Aniketh Garikaparthi, Manasi Patwardhan +1

Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on weak update mechanisms, such as full-prompt…

cs.CL2025

Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights

Hyunjae Kim, Jiwoong Sohn, Aidan Gilson +24

Large language models (LLMs) are transforming the landscape of medicine, yet two fundamental challenges persist: keeping up with rapidly evolving medical knowledge and providing ve…

cs.CL2024

Bayesian Calibration of Win Rate Estimation with LLM Evaluators

Yicheng Gao, Gonghan Xu, Zhe Wang +1

Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evalua…