Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
arXiv:2007.15779 · doi:10.1145/3458754
Abstract
Pretraining large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. However, most pretraining efforts focus on general domain corpora, such as newswire and Web. A prevailing assumption is that even domain-specific pretraining can benefit by starting from general-domain language models. In this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining of general-domain language models. To facilitate this investigation, we compile a comprehensive biomedical NLP benchmark from publicly-available datasets. Our experiments show that domain-specific pretraining serves as a solid foundation for a wide range of biomedical NLP tasks, leading to new state-of-the-art results across the board. Further, in conducting a thorough evaluation of modeling choices, both for pretraining and task-specific fine-tuning, we discover that some common practices are unnecessary with BERT models, such as using complex tagging schemes in named entity recognition (NER). To help accelerate research in biomedical NLP, we have released our state-of-the-art pretrained and task-specific models for the community, and created a leaderboard featuring our BLURB benchmark (short for Biomedical Language Understanding & Reasoning Benchmark) at https://aka.ms/BLURB.
ACM Transactions on Computing for Healthcare (HEALTH)
References in corpus (1)
Cited by in corpus (153)
- BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining
- Deep Learning Based Text Classification: A Comprehensive Review
- Large language models in medicine: the potentials and pitfalls
- ChatGPT for Shaping the Future of Dentistry: The Potential of Multi-Modal Large Language Model
- Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing
- Matching Patients to Clinical Trials with Large Language Models
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Drug Repurposing for COVID-19 via Knowledge Graph Completion
- Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
- Transformers in Healthcare: A Survey
- Benchmarking large language models for biomedical natural language processing applications and recommendations
- MedCPT: Contrastive Pre-trained Transformers with Large-scale PubMed Search Logs for Zero-shot Biomedical Information Retrieval
- BioRED: A Rich Biomedical Relation Extraction Dataset
- Improving Chest X-Ray Report Generation by Leveraging Warm Starting
- A general-purpose material property data extraction pipeline from large polymer corpora using Natural Language Processing
- Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary
- BERN2: an advanced neural biomedical named entity recognition and normalization tool
- Biomedical Question Answering: A Survey of Approaches and Challenges
- Pretrained Transformers for Text Ranking: BERT and Beyond
- BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once
- A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine
- Does the Magic of BERT Apply to Medical Code Assignment? A Quantitative Study
- CLIP in Medical Imaging: A Survey
- Automated Radiology Report Generation: A Review of Recent Advances
- AIONER: All-in-one scheme-based biomedical named entity recognition using deep learning
- Mining Legal Arguments in Court Decisions
- A Large Language Model Approach to Educational Survey Feedback Analysis
- GPT-4 can pass the Korean National Licensing Examination for Korean Medicine Doctors
- Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study
- MEDBERT.de: A Comprehensive German BERT Model for the Medical Domain
- Enhancing Knowledge Retrieval with In-Context Learning and Semantic Search through Generative AI
- ThoughtSource: A central hub for large language model reasoning data
- Hierarchical Label-wise Attention Transformer Model for Explainable ICD Coding
- Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
- Ten Quick Tips for Deep Learning in Biology
- Comparison of biomedical relationship extraction methods and models for knowledge graph creation
- ERNIE-GeoL: A Geography-and-Language Pre-trained Model and its Applications in Baidu Maps
- Artificial General Intelligence for Medical Imaging Analysis
- Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
- Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- PMC-Patients: A Large-scale Dataset of Patient Summaries and Relations for Benchmarking Retrieval-based Clinical Decision Support Systems
- De-identification of clinical free text using natural language processing: A systematic review of current approaches
- Correctness Comparison of ChatGPT-4, Gemini, Claude-3, and Copilot for Spatial Tasks
- Automated ICD Coding using Extreme Multi-label Long Text Transformer-based Models
- nach0: Multimodal Natural and Chemical Languages Foundation Model
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- Several categories of Large Language Models (LLMs): A Short Survey
- The Future of Intelligent Healthcare: A Systematic Analysis and Discussion on the Integration and Impact of Robots Using Large Language Models for Healthcare
- Clinical Insights: A Comprehensive Review of Language Models in Medicine
- How Do Your Biomedical Named Entity Recognition Models Generalize to Novel Entities?
- Ontology-Driven and Weakly Supervised Rare Disease Identification from Clinical Notes
- Informing clinical assessment by contextualizing post-hoc explanations of risk prediction models in type-2 diabetes
- Bridging the Gap between Chemical Reaction Pretraining and Conditional Molecule Generation with a Unified Model
- Context Matters: A Strategy to Pre-train Language Model for Science Education
- Learning to Automate Follow-up Question Generation using Process Knowledge for Depression Triage on Reddit Posts
- Evaluating Pre-trained Convolutional Neural Networks and Foundation Models as Feature Extractors for Content-based Medical Image Retrieval
- Sequence tagging for biomedical extractive question answering
- DR.BENCH: Diagnostic Reasoning Benchmark for Clinical Natural Language Processing
- Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario
- Data-Centric Foundation Models in Computational Healthcare: A Survey
- KnowledgeVIS: Interpreting Language Models by Comparing Fill-in-the-Blank Prompts
- Evaluation of GPT and BERT-based models on identifying protein-protein interactions in biomedical text
- Read, Attend, and Code: Pushing the Limits of Medical Codes Prediction from Clinical Notes by Machines
- From Zero to Hero: Harnessing Transformers for Biomedical Named Entity Recognition in Zero- and Few-shot Contexts
- Neural Rankers for Effective Screening Prioritisation in Medical Systematic Review Literature Search
- Pre-training technique to localize medical BERT and enhance biomedical BERT
- Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
- Extracting Medication Changes in Clinical Narratives using Pre-trained Language Models
- ELECTRAMed: a new pre-trained language representation model for biomedical NLP
- OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining
- Pre-Trained Language Models for Keyphrase Prediction: A Review
- Scientific Language Models for Biomedical Knowledge Base Completion: An Empirical Study
- Self-supervised learning-based general laboratory progress pretrained model for cardiovascular event detection
- Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking
- The Leaf Clinical Trials Corpus: a new resource for query generation from clinical trial eligibility criteria
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- Large-vocabulary forensic pathological analyses via prototypical cross-modal contrastive learning
- Domain-Specific Pretraining for Vertical Search: Case Study on Biomedical Literature
- Toward expanding the scope of radiology report summarization to multiple anatomies and modalities
- Leveraging Large Language Models for Knowledge-free Weak Supervision in Clinical Natural Language Processing
- A Meta-Evaluation of Faithfulness Metrics for Long-Form Hospital-Course Summarization
- Building Chinese Biomedical Language Models via Multi-Level Text Discrimination
- DyGen: Learning from Noisy Labels via Dynamics-Enhanced Generative Modeling
- CCS Explorer: Relevance Prediction, Extractive Summarization, and Named Entity Recognition from Clinical Cohort Studies
- DiMB-RE: Mining the Scientific Literature for Diet-Microbiome Associations
- Improving Biomedical Pretrained Language Models with Knowledge
- Medical Coding with Biomedical Transformer Ensembles and Zero/Few-shot Learning
- POPDx: An Automated Framework for Patient Phenotyping across 392,246 Individuals in the UK Biobank Study
- Advancing Italian Biomedical Information Extraction with Transformers-based Models: Methodological Insights and Multicenter Practical Application
- HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights
- Overview of BioASQ 2024: The twelfth BioASQ challenge on Large-Scale Biomedical Semantic Indexing and Question Answering
- Natural Language-Assisted Multi-modal Medication Recommendation
- Extracting Protein-Protein Interactions (PPIs) from Biomedical Literature using Attention-based Relational Context Information
- Reviewing Clinical Knowledge in Medical Large Language Models: Training and Beyond
- A Generalizable Deep Learning System for Cardiac MRI
- Augmenting Biomedical Named Entity Recognition with General-domain Resources
- ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering
- A Biomedical Entity Extraction Pipeline for Oncology Health Records in Portuguese
- A Hybrid Approach to Measure Semantic Relatedness in Biomedical Concepts
- SMILE: Evaluation and Domain Adaptation for Social Media Language Understanding
- Integrating Chain-of-Thought and Retrieval Augmented Generation Enhances Rare Disease Diagnosis from Clinical Notes
- Improving Tagging Consistency and Entity Coverage for Chemical Identification in Full-text Articles
- Assigning Species Information to Corresponding Genes by a Sequence Labeling Framework
- Extracting Lifestyle Factors for Alzheimer's Disease from Clinical Notes Using Deep Learning with Weak Supervision
- Reshaping MOFs text mining with a dynamic multi-agents framework of large language model
- Improving BERT Model Using Contrastive Learning for Biomedical Relation Extraction
- Large-scale investigation of weakly-supervised deep learning for the fine-grained semantic indexing of biomedical literature
- Biomedical Entity Linking with Contrastive Context Matching
- Knowledge-augmented Pre-trained Language Models for Biomedical Relation Extraction
- GOProteinGNN: Leveraging Protein Knowledge Graphs for Protein Representation Learning
- Combining Evidence and Reasoning for Biomedical Fact-Checking
- Dense Retrieval with Continuous Explicit Feedback for Systematic Review Screening Prioritisation
- Contextual embedding and model weighting by fusing domain knowledge on Biomedical Question Answering
- Variational Open-Domain Question Answering
- CancerBERT: a BERT model for Extracting Breast Cancer Phenotypes from Electronic Health Records
- An Experimental Evaluation of Transformer-based Language Models in the Biomedical Domain
- Enhancing Biomedical Knowledge Discovery for Diseases: An Open-Source Framework Applied on Rett Syndrome and Alzheimer's Disease
- Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
- A New Entity Extraction Method Based on Machine Reading Comprehension
- Deep learning models are not robust against noise in clinical text
- Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation
- Leveraging Knowledge Graph Embeddings to Enhance Contextual Representations for Relation Extraction
- SHADE: Semantic Hypernym Annotator for Domain-specific Entities -- DnD Domain Use Case
- BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
- Predicting COVID-19 Patient Shielding: A Comprehensive Study
- Pre-training Language Model Incorporating Domain-specific Heterogeneous Knowledge into A Unified Representation
- AF Adapter: Continual Pretraining for Building Chinese Biomedical Language Model
- Analyzing Research Trends in Inorganic Materials Literature Using NLP
- Search-Optimized Quantization in Biomedical Ontology Alignment
- Large Language Models as AI Agents for Digital Atoms and Molecules: Catalyzing a New Era in Computational Biophysics
- A Library Perspective on Supervised Text Processing in Digital Libraries: An Investigation in the Biomedical Domain
- A reproducible experimental survey on biomedical sentence similarity: a string-based method sets the state of the art
- Data Overdose? Time for a Quadruple Shot: Knowledge Graph Construction using Enhanced Triple Extraction
- Supporting Evidence-Based Medicine by Finding Both Relevant and Significant Works
- BCH-NLP at BioCreative VII Track 3: medications detection in tweets using transformer networks and multi-task learning
- SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text Mining
- Recent Advances in Automated Question Answering In Biomedical Domain
- How Important is Domain Specificity in Language Models and Instruction Finetuning for Biomedical Relation Extraction?
- BERT might be Overkill: A Tiny but Effective Biomedical Entity Linker based on Residual Convolutional Neural Networks
- Multimodal Integrated Knowledge Transfer to Large Language Models through Preference Optimization with Biomedical Applications
- Generating multiple-choice questions for medical question answering with distractors and cue-masking
- Searching Personal Collections
- Exploring a Unified Sequence-To-Sequence Transformer for Medical Product Safety Monitoring in Social Media
- Exploring Low-Cost Transformer Model Compression for Large-Scale Commercial Reply Suggestions
- BagBERT: BERT-based bagging-stacking for multi-topic classification
- On the Universality of Deep Contextual Language Models
- Description-based Label Attention Classifier for Explainable ICD-9 Classification
- Using Text Analytics for Health to Get Meaningful Insights from a Corpus of COVID Scientific Papers
- ClusterChat: Multi-Feature Search for Corpus Exploration
- Discovering novel drug-supplement interactions using a dietary supplements knowledge graph generated from the biomedical literature