BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
arXiv:2210.10341 · doi:10.1093/bib/bbac409
Abstract
Pre-trained language models have attracted increasing attention in the biomedical domain, inspired by their great success in the general natural language domain. Among the two main branches of pre-trained language models in the general language domain, i.e., BERT (and its variants) and GPT (and its variants), the first one has been extensively studied in the biomedical domain, such as BioBERT and PubMedBERT. While they have achieved great success on a variety of discriminative downstream biomedical tasks, the lack of generation ability constrains their application scope. In this paper, we propose BioGPT, a domain-specific generative Transformer language model pre-trained on large scale biomedical literature. We evaluate BioGPT on six biomedical NLP tasks and demonstrate that our model outperforms previous models on most tasks. Especially, we get 44.98%, 38.42% and 40.76% F1 score on BC5CDR, KD-DTI and DDI end-to-end relation extraction tasks respectively, and 78.2% accuracy on PubMedQA, creating a new record. Our case study on text generation further demonstrates the advantage of BioGPT on biomedical literature to generate fluent descriptions for biomedical terms. Code is available at https://github.com/microsoft/BioGPT.
Published at Briefings in Bioinformatics. Code is available at https://github.com/microsoft/BioGPT
References in corpus (4)
- Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing
- SciFive: a text-to-text transformer model for biomedical literature
- ELECTRAMed: a new pre-trained language representation model for biomedical NLP
- A sequence-to-sequence approach for document-level relation extraction
Cited by in corpus (48)
- A Study of Generative Large Language Model for Medical Research and Healthcare
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Large language models in medicine: the potentials and pitfalls
- BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Transformers in Healthcare: A Survey
- GeneGPT: Augmenting Large Language Models with Domain Tools for Improved Access to Biomedical Information
- Benchmarking large language models for biomedical natural language processing applications and recommendations
- Prompt engineering paradigms for medical applications: scoping review and recommendations for better practices
- Structured prompt interrogation and recursive extraction of semantics (SPIRES): A method for populating knowledge bases using zero-shot learning
- Current and future directions in network biology
- GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
- Taiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical Tasks
- Differentiate ChatGPT-generated and Human-written Medical Texts
- Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models
- MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
- From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
- The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
- nach0: Multimodal Natural and Chemical Languages Foundation Model
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- MedSyn: LLM-based Synthetic Medical Text Generation Framework
- Evaluating Pre-trained Convolutional Neural Networks and Foundation Models as Feature Extractors for Content-based Medical Image Retrieval
- Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
- From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
- Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation
- Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models
- A survey on cutting-edge relation extraction techniques based on language models
- Overview of the BioLaySumm 2023 Shared Task on Lay Summarization of Biomedical Research Articles
- Leveraging Large Language Models through Natural Language Processing to provide interpretable Machine Learning predictions of mental deterioration in real time
- Enhancing chest X-ray datasets with privacy-preserving large language models and multi-type annotations: a data-driven approach for improved classification
- CURENet: Combining Unified Representations for Efficient Chronic Disease Prediction
- Augmenting Biomedical Named Entity Recognition with General-domain Resources
- LOCO-EPI: Leave-one-chromosome-out (LOCO) as a benchmarking paradigm for deep learning based prediction of enhancer-promoter interactions
- SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
- Generative Language Models on Nucleotide Sequences of Human Genes
- Predicting drug-gene relations via analogy tasks with word embeddings
- Combining Evidence and Reasoning for Biomedical Fact-Checking
- MIRAGE: Retrieval and Generation of Multimodal Images and Texts for Medical Education
- Automated Spinal MRI Labelling from Reports Using a Large Language Model
- A Survey on Memory-Efficient Transformer-Based Model Training in AI for Science
- Enhancing Biomedical Knowledge Discovery for Diseases: An Open-Source Framework Applied on Rett Syndrome and Alzheimer's Disease
- Large Language Models as AI Agents for Digital Atoms and Molecules: Catalyzing a New Era in Computational Biophysics
- How Important is Domain Specificity in Language Models and Instruction Finetuning for Biomedical Relation Extraction?
- Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
- FinBloom: Knowledge Grounding Large Language Model with Real-time Financial Data
- Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation: A Comparative Study
- Comparison of pipeline, sequence-to-sequence, and GPT models for end-to-end relation extraction: experiments with the rare disease use-case