Clinical Insights: A Comprehensive Review of Language Models in Medicine
arXiv:2408.11735 · doi:10.1371/journal.pdig.0000800
Abstract
This paper explores the advancements and applications of language models in healthcare, focusing on their clinical use cases. It examines the evolution from early encoder-based systems requiring extensive fine-tuning to state-of-the-art large language and multimodal models capable of integrating text and visual data through in-context learning. The analysis emphasizes locally deployable models, which enhance data privacy and operational autonomy, and their applications in tasks such as text generation, classification, information extraction, and conversational systems. The paper also highlights a structured organization of tasks and a tiered ethical approach, providing a valuable resource for researchers and practitioners, while discussing key challenges related to ethics, evaluation, and implementation.
Submitted to PLOS Digital Health, Revision 1
References in corpus (62)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Learning Transferable Visual Models From Natural Language Supervision
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- On the Opportunities and Risks of Foundation Models
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Scaling Instruction-Finetuned Language Models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Gemini: A Family of Highly Capable Multimodal Models
- Visual Instruction Tuning
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- BARTScore: Evaluating Generated Text as Text Generation
- Measuring Massive Multitask Language Understanding
- Gemma: Open Models Based on Gemini Research and Technology
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Towards Interpretable Mental Health Analysis with Large Language Models
- The Falcon Series of Open Language Models
- MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
- BIMCV COVID-19+: a large annotated dataset of RX and CT images from COVID-19 patients
- UL2: Unifying Language Learning Paradigms
- SciFive: a text-to-text transformer model for biomedical literature
- Dissociating language and thought in large language models
- DeID-GPT: Zero-shot Medical Text De-Identification by GPT-4
- MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
- LLMs Accelerate Annotation for Medical Information Extraction
- Bias and Fairness in Large Language Models: A Survey
- PMC-LLaMA: Towards Building Open-source Language Models for Medicine
- GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
- Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
- Conversational Health Agents: A Personalized LLM-Powered Agent Framework
- A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
- Explainable Natural Language Processing for Corporate Sustainability Analysis
- How Interpretable are Reasoning Explanations from Prompting Large Language Models?
- Human-AI Collaboration Enables More Empathic Conversations in Text-based Peer-to-Peer Mental Health Support
- LAVIS: A Library for Language-Vision Intelligence
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Analyzing Leakage of Personally Identifiable Information in Language Models
- PaliGemma: A versatile 3B VLM for transfer
- PharmacyGPT: The AI Pharmacist
- An Empirical Evaluation of Prompting Strategies for Large Language Models in Zero-Shot Clinical Natural Language Processing
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- Dia-LLaMA: Towards Large Language Model-driven CT Report Generation
- DynaEval: Unifying Turn and Dialogue Level Evaluation
- DePlot: One-shot visual language reasoning by plot-to-table translation
- Cognitive Reframing of Negative Thoughts through Human-Language Model Interaction
- PreCog: Exploring the Relation between Memorization and Performance in Pre-trained Language Models
- Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models
- Investigating Efficiently Extending Transformers for Long Input Summarization
- An Annotated Dataset for Explainable Interpersonal Risk Factors of Mental Disturbance in Social Media Posts
- Revisiting text decomposition methods for NLI-based factuality scoring of summaries
- Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models
- GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
- Interpretable Differential Diagnosis with Dual-Inference Large Language Models
- Elephants Never Forget: Testing Language Models for Memorization of Tabular Data
- An evaluation of GPT models for phenotype concept recognition
- Comparing Two Model Designs for Clinical Note Generation; Is an LLM a Useful Evaluator of Consistency?