Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices
arXiv:2405.01249 · doi:10.2196/60501
Abstract
Prompt engineering is crucial for harnessing the potential of large language models (LLMs), especially in the medical domain where specialized terminology and phrasing is used. However, the efficacy of prompt engineering in the medical domain remains to be explored. In this work, 114 recent studies (2022-2024) applying prompt engineering in medicine, covering prompt learning (PL), prompt tuning (PT), and prompt design (PD) are reviewed. PD is the most prevalent (78 articles). In 12 papers, PD, PL, and PT terms were used interchangeably. ChatGPT is the most commonly used LLM, with seven papers using it for processing sensitive clinical data. Chain-of-Thought emerges as the most common prompt engineering technique. While PL and PT articles typically provide a baseline for evaluating prompt-based approaches, 64% of PD studies lack non-prompt-related baselines. We provide tables and figures summarizing existing work, and reporting recommendations to guide future research contributions.
References in corpus (26)
- A Survey of Large Language Models
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Capabilities of GPT-4 on Medical Challenge Problems
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data
- Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
- Towards Interpretable Mental Health Analysis with Large Language Models
- Artificial Intelligence for Health Message Generation: Theory, Method, and an Empirical Study Using Prompt Engineering
- Enhancing Phenotype Recognition in Clinical Notes Using Large Language Models: PhenoBCBERT and PhenoGPT
- Model Tuning or Prompt Tuning? A Study of Large Language Models for Clinical Concept and Relation Extraction
- OpenPrompt: An Open-source Framework for Prompt-learning
- An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
- LLM-empowered Chatbots for Psychiatrist and Patient Simulation: Application and Evaluation
- HealthPrompt: A Zero-shot Learning Paradigm for Clinical Natural Language Processing
- Automated Identification of Eviction Status from Electronic Health Record Notes
- CliniDigest: A Case Study in Large Language Model Based Large-Scale Summarization of Clinical Trial Descriptions
- Position: Key Claims in LLM Research Have a Long Tail of Footnotes
- From Beginner to Expert: Modeling Medical Knowledge into General LLMs
- Keyword-optimized Template Insertion for Clinical Information Extraction via Prompt-based Learning
- Investigating Large Language Models and Control Mechanisms to Improve Text Readability of Biomedical Abstracts
- Clinical Text Deduplication Practices for Efficient Pretraining and Improved Clinical Tasks
- Large Language Models and Prompt Engineering for Biomedical Query Focused Multi-Document Summarisation
- Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
- Overview of the PromptCBLUE Shared Task in CHIP2023
- ChatGPT and post-test probability