4 papers · 1 filter
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
Asma Ben Abacha, Wen-wai Yim, Yujuan Fu +4
Several studies showed that Large Language Models (LLMs) can answer medical questions correctly, even outperforming the average human score in some medical exams. However, to our k…
Find Parent then Label Children: A Two-stage Taxonomy Completion Method with Pre-trained Language Model
Fei Xia, Yixuan Weng, Shizhu He +2
Taxonomies, which organize domain concepts into hierarchical structures, are crucial for building knowledge systems and downstream applications. As domain knowledge evolves, taxono…
Large Language Models Need Holistically Thought in Medical Conversational QA
Yixuan Weng, Bin Li, Fei Xia +5
The medical conversational question answering (CQA) system aims at providing a series of professional medical services to improve the efficiency of medical care. Despite the succes…
NLPStatTest: A Toolkit for Comparing NLP System Performance
Haotian Zhu, Denise Mak, Jesse Gioannini +1
Statistical significance testing centered on p-values is commonly used to compare NLP system performance, but p-values alone are insufficient because statistical significance diffe…