MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
arXiv:2404.05590 · doi:10.1016/j.artmed.2024.102938
Abstract
Large Language Models (LLMs) have the potential of facilitating the development of Artificial Intelligence technology to assist medical experts for interactive decision support, which has been demonstrated by their competitive performances in Medical QA. However, while impressive, the required quality bar for medical applications remains far from being achieved. Currently, LLMs remain challenged by outdated knowledge and by their tendency to generate hallucinated content. Furthermore, most benchmarks to assess medical knowledge lack reference gold explanations which means that it is not possible to evaluate the reasoning of LLMs predictions. Finally, the situation is particularly grim if we consider benchmarking LLMs for languages other than English which remains, as far as we know, a totally neglected topic. In order to address these shortcomings, in this paper we present MedExpQA, the first multilingual benchmark based on medical exams to evaluate LLMs in Medical Question Answering. To the best of our knowledge, MedExpQA includes for the first time reference gold explanations written by medical doctors which can be leveraged to establish various gold-based upper-bounds for comparison with LLMs performance. Comprehensive multilingual experimentation using both the gold reference explanations and Retrieval Augmented Generation (RAG) approaches show that performance of LLMs still has large room for improvement, especially for languages other than English. Furthermore, and despite using state-of-the-art RAG methods, our results also demonstrate the difficulty of obtaining and integrating readily available medical knowledge that may positively impact results on downstream evaluations for Medical Question Answering. So far the benchmark is available in four languages, but we hope that this work may encourage further development to other languages.
References in corpus (13)
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- PaLM: Scaling Language Modeling with Pathways
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Capabilities of GPT-4 on Medical Challenge Problems
- Towards Expert-Level Medical Question Answering with Large Language Models
- MedCPT: Contrastive Pre-trained Transformers with Large-scale PubMed Search Logs for Zero-shot Biomedical Information Retrieval
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- SciFive: a text-to-text transformer model for biomedical literature
- PMC-LLaMA: Towards Building Open-source Language Models for Medicine
- ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
- Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
- HiTZ@Antidote: Argumentation-driven Explainable Artificial Intelligence for Digital Medicine