Benchmarking large language models for biomedical natural language processing applications and recommendations
arXiv:2305.16326 · doi:10.1038/s41467-025-56989-2
Abstract
The rapid growth of biomedical literature poses challenges for manual knowledge curation and synthesis. Biomedical Natural Language Processing (BioNLP) automates the process. While Large Language Models (LLMs) have shown promise in general domains, their effectiveness in BioNLP tasks remains unclear due to limited benchmarks and practical guidelines. We perform a systematic evaluation of four LLMs, GPT and LLaMA representatives on 12 BioNLP benchmarks across six applications. We compare their zero-shot, few-shot, and fine-tuning performance with traditional fine-tuning of BERT or BART models. We examine inconsistencies, missing information, hallucinations, and perform cost analysis. Here we show that traditional fine-tuning outperforms zero or few shot LLMs in most tasks. However, closed-source LLMs like GPT-4 excel in reasoning-related tasks such as medical question answering. Open source LLMs still require fine-tuning to close performance gaps. We find issues like missing information and hallucinations in LLM outputs. These results offer practical insights for applying LLMs in BioNLP.
References in corpus (16)
- SciPy 1.0--Fundamental Algorithms for Scientific Computing in Python
- BioBERT: a pre-trained biomedical language representation model for biomedical text mining
- Training language models to follow instructions with human feedback
- Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Large language models in medicine: the potentials and pitfalls
- ML-Net: multi-label classification of biomedical texts with deep neural networks
- BioSentVec: creating sentence embeddings for biomedical texts
- Biomedical Question Answering: A Survey of Approaches and Challenges
- An Empirical Survey on Long Document Summarization: Datasets, Models and Metrics
- Deep Bidirectional Language-Knowledge Graph Pretraining
- Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
- BioConceptVec: creating and evaluating literature-based biomedical concept embeddings on a large scale
- Scaling Clinical Trial Matching Using Large Language Models: A Case Study in Oncology
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning