most citedMEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

9 citations · 9 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review

Zhaoyi Sun, Minal Jagtiani, Wen-wai Yim +4

Radiology reports for the same patient examination may contain clinically meaningful discrepancies arising from interpretation differences, reporting variability, or evolving asses…

cs.CL2025

Automated Identification of Incidentalomas Requiring Follow-Up: A Multi-Anatomy Evaluation of LLM-Based and Supervised Approaches

Namu Park, Farzad Ahmed, Zhaoyi Sun +6

Objective: To evaluate large language models (LLMs) against supervised baselines for fine-grained, lesion-level detection of incidentalomas requiring follow-up, addressing the limi…

cs.CL2025

UW-BioNLP at ChemoTimelines 2025: Thinking, Fine-Tuning, and Dictionary-Enhanced LLM Systems for Chemotherapy Timeline Extraction

Tianmai M. Zhang, Zhaoyi Sun, Sihang Zeng +5

The ChemoTimelines shared task benchmarks methods for constructing timelines of systemic anticancer treatment from electronic health records of cancer patients. This paper describe…

cs.CL2025

A Scoping Review of Natural Language Processing in Addressing Medically Inaccurate Information: Errors, Misinformation, and Hallucination

Zhaoyi Sun, Wen-Wai Yim, Ozlem Uzuner +2

Objective: This review aims to explore the potential and challenges of using Natural Language Processing (NLP) to detect, correct, and mitigate medically inaccurate information, in…

cs.CL20259 cited

MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes

Asma Ben Abacha, Wen-wai Yim, Yujuan Fu +4

Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our k…