Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
arXiv:2306.10070 · doi:10.1093/bib/bbad493
Abstract
ChatGPT has drawn considerable attention from both the general public and domain experts with its remarkable text generation capabilities. This has subsequently led to the emergence of diverse applications in the field of biomedicine and health. In this work, we examine the diverse applications of large language models (LLMs), such as ChatGPT, in biomedicine and health. Specifically we explore the areas of biomedical information retrieval, question answering, medical text summarization, information extraction, and medical education, and investigate whether LLMs possess the transformative power to revolutionize these tasks or whether the distinct complexities of biomedical domain presents unique challenges. Following an extensive literature survey, we find that significant advances have been made in the field of text generation tasks, surpassing the previous state-of-the-art methods. For other applications, the advances have been modest. Overall, LLMs have not yet revolutionized biomedicine, but recent rapid progress indicates that such methods hold great potential to provide valuable means for accelerating discovery and improving health. We also find that the use of LLMs, like ChatGPT, in the fields of biomedicine and health entails various risks and challenges, including fabricated information in its generated responses, as well as legal and privacy concerns associated with sensitive patient data. We believe this survey can provide a comprehensive and timely overview to biomedical researchers and healthcare practitioners on the opportunities and challenges associated with using ChatGPT and other LLMs for transforming biomedicine and health.
References in corpus (19)
- Training language models to follow instructions with human feedback
- A Survey of Large Language Models
- Scaling Instruction-Finetuned Language Models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Capabilities of GPT-4 on Medical Challenge Problems
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Towards Expert-Level Medical Question Answering with Large Language Models
- Matching Patients to Clinical Trials with Large Language Models
- GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
- Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
- Structured prompt interrogation and recursive extraction of semantics (SPIRES): A method for populating knowledge bases using zero-shot learning
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- PAL: Program-aided Language Models
- HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
- Deep Bidirectional Language-Knowledge Graph Pretraining
- AIONER: All-in-one scheme-based biomedical named entity recognition using deep learning
- DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
- TALM: Tool Augmented Language Models
- ELECTRAMed: a new pre-trained language representation model for biomedical NLP
Cited by in corpus (24)
- Matching Patients to Clinical Trials with Large Language Models
- GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
- Benchmarking large language models for biomedical natural language processing applications and recommendations
- Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
- Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices
- PubMed and Beyond: Biomedical Literature Search in the Age of Artificial Intelligence
- Redefining Qualitative Analysis in the AI Era: Utilizing ChatGPT for Efficient Thematic Analysis
- Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation
- Taiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical Tasks
- Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study
- Understanding Users' Dissatisfaction with ChatGPT Responses: Types, Resolving Tactics, and the Effect of Knowledge Level
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- The current status of large language models in summarizing radiology report impressions
- Mapping the individual, social, and biospheric impacts of Foundation Models
- Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning
- A Simplified Retriever to Improve Accuracy of Phenotype Normalizations by Large Language Models
- Unify Graph Learning with Text: Unleashing LLM Potentials for Session Search
- Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports
- CoGrader: Transforming Instructors' Assessment of Project Reports through Collaborative LLM Integration
- Cancer Diagnosis Categorization in Electronic Health Records Using Large Language Models and BioBERT: Model Performance Evaluation Study
- Enhancing Biomedical Knowledge Discovery for Diseases: An Open-Source Framework Applied on Rett Syndrome and Alzheimer's Disease
- Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation
- Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation: A Comparative Study
- Evaluation of the phi-3-mini SLM for identification of texts related to medicine, health, and sports injuries