A Study of Generative Large Language Model for Medical Research and Healthcare
arXiv:2305.13523 · doi:10.1038/s41746-023-00958-w
Abstract
There is enormous enthusiasm and concerns in using large language models (LLMs) in healthcare, yet current assumptions are all based on general-purpose LLMs such as ChatGPT. This study develops a clinical generative LLM, GatorTronGPT, using 277 billion words of mixed clinical and English text with a GPT-3 architecture of 20 billion parameters. GatorTronGPT improves biomedical natural language processing for medical research. Synthetic NLP models trained using GatorTronGPT generated text outperform NLP models trained using real-world clinical text. Physicians Turing test using 1 (worst) to 9 (best) scale shows that there is no significant difference in linguistic readability (p = 0.22; 6.57 of GatorTronGPT compared with 6.93 of human) and clinical relevance (p = 0.91; 7.0 of GatorTronGPT compared with 6.97 of human) and that physicians cannot differentiate them (p < 0.001). This study provides insights on the opportunities and challenges of LLMs for medical research and healthcare.
References in corpus (3)
Cited by in corpus (14)
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
- MedSyn: LLM-based Synthetic Medical Text Generation Framework
- Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE Questions
- Clinical information extraction for Low-resource languages with Few-shot learning using Pre-trained language models and Prompting
- A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectives
- Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
- Robust Privacy Amidst Innovation with Large Language Models Through a Critical Assessment of the Risks
- DualSG: A Dual-Stream Explicit Semantic-Guided Multivariate Time Series Forecasting Framework
- Generative AI in Health Economics and Outcomes Research: A Taxonomy of Key Definitions and Emerging Applications, an ISPOR Working Group Report
- Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
- Automatic Summarization of Doctor-Patient Encounter Dialogues Using Large Language Model through Prompt Tuning
- Mapping Scientific Literature with Large Language Models and Topic Modeling
- SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models