Advancing bioinformatics with large language models: components, applications and perspectives
arXiv:2401.04155 · doi:10.1093/bib/bbag367
Abstract
Large language models (LLMs) are a class of artificial intelligence models based on deep learning, which have great performance in various tasks, especially in natural language processing (NLP). Large language models typically consist of artificial neural networks with numerous parameters, trained on large amounts of unlabeled input using self-supervised or semi-supervised learning. However, their potential for solving bioinformatics problems may even exceed their proficiency in modeling human language. In this review, we will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics, spanning genomics, transcriptomics, proteomics, drug discovery, and single-cell analysis. Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, the core attention mechanism, and the pre-training processes underlying these models. Additionally, we will introduce currently available foundation models and highlight their downstream applications across various bioinformatics domains. Finally, drawing from our experience, we will offer practical guidance for both LLM users and developers, emphasizing strategies to optimize their use and foster further innovation in the field.
5 main figures
References in corpus (18)
- Distilling the Knowledge in a Neural Network
- ProtVec: A Continuous Distributed Representation of Biological Sequences
- Protein-Ligand Scoring with Convolutional Neural Networks
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- Controllable Protein Design with Language Models
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome
- RiNALMo: General-Purpose RNA Language Models Can Generalize Well on Structure Prediction Tasks
- Interpretable RNA Foundation Model from Unannotated Data for Highly Accurate RNA Structure and Function Predictions
- OntoProtein: Protein Pretraining With Gene Ontology Embedding
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical Texts
- Single-Cell Multimodal Prediction via Transformers
- Large-Scale Cell Representation Learning via Divide-and-Conquer Contrastive Learning
- Single Cells Are Spatial Tokens: Transformers for Spatial Transcriptomic Data Imputation
- On Pre-trained Language Models for Antibody
- DrugCLIP: Contrastive Drug-Disease Interaction For Drug Repurposing
- Protein Representation Learning via Knowledge Enhanced Primary Structure Modeling
- RNA-GPT: Multimodal Generative System for RNA Sequence Understanding
- A Comparative Review of RNA Language Models