3 papers
cs.CL2026
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
Zhimeng Luo, Lixin Wu, Adam Frisch +1
Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally info…
cs.CL2025
Weakly Supervised Medical Entity Extraction and Linking for Chief Complaints
Zhimeng Luo, Zhendong Wang, Rui Meng +3
A Chief complaint (CC) is the reason for the medical visit as stated in the patient's own words. It helps medical professionals to quickly understand a patient's situation, and als…
cs.CL2025
Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning
Zhimeng Luo, Abhibha Gupta, Adam Frisch +1
The extraction of critical patient information from Electronic Health Records (EHRs) poses significant challenges due to the complexity and unstructured nature of the data. Traditi…