SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery
arXiv:1911.04738
Abstract
In drug-discovery-related tasks such as virtual screening, machine learning is emerging as a promising way to predict molecular properties. Conventionally, molecular fingerprints (numerical representations of molecules) are calculated through rule-based algorithms that map molecules to a sparse discrete space. However, these algorithms perform poorly for shallow prediction models or small datasets. To address this issue, we present SMILES Transformer. Inspired by Transformer and pre-trained language models from natural language processing, SMILES Transformer learns molecular fingerprints through unsupervised pre-training of the sequence-to-sequence language model using a huge corpus of SMILES, a text representation system for molecules. We performed benchmarks on 10 datasets against existing fingerprints and graph-based methods and demonstrated the superiority of the proposed algorithms in small-data settings where pre-training facilitated good generalization. Moreover, we define a novel metric to concurrently measure model accuracy and data efficiency.
9 pages, 4 figures
References in corpus (2)
Cited by in corpus (15)
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- A smile is all you need: Predicting limiting activity coefficients from SMILES with natural language processing
- SPT-NRTL: A physics-guided machine learning model to predict thermodynamically consistent activity coefficients
- Empowering Molecule Discovery for Molecule-Caption Translation with Large Language Models: A ChatGPT Perspective
- Chemical-Reaction-Aware Molecule Representation Learning
- Artificial Intelligence in Materials Science and Engineering: Current Landscape, Key Challenges, and Future Trajectorie
- Multimodal Model with Text and Drug Embeddings for Adverse Drug Reaction Classification
- Learn molecular representations from large-scale unlabeled molecules for drug discovery
- Dual-view Molecule Pre-training
- BatmanNet: Bi-branch Masked Graph Transformer Autoencoder for Molecular Representation
- Learning Attributed Graph Representations with Communicative Message Passing Transformer
- Artificial Intelligence in Drug Discovery: Applications and Techniques
- A Bayesian Flow Network Framework for Chemistry Tasks
- Distributed Deep Learning in Open Collaborations
- Using Clinical Drug Representations for Improving Mortality and Length of Stay Predictions