SMILES Enumeration as Data Augmentation for Neural Network Modeling of Molecules
arXiv:1703.07076
Abstract
Simplified Molecular Input Line Entry System (SMILES) is a single line text representation of a unique molecule. One molecule can however have multiple SMILES strings, which is a reason that canonical SMILES have been defined, which ensures a one to one correspondence between SMILES string and molecule. Here the fact that multiple SMILES represent the same molecule is explored as a technique for data augmentation of a molecular QSAR dataset modeled by a long short term memory (LSTM) cell based neural network. The augmented dataset was 130 times bigger than the original. The network trained with the augmented dataset shows better performance on a test set when compared to a model built with only one canonical SMILES string per molecule. The correlation coefficient R2 on the test set was improved from 0.56 to 0.66 when using SMILES enumeration, and the root mean square error (RMS) likewise fell from 0.62 to 0.55. The technique also works in the prediction phase. By taking the average per molecule of the predictions for the enumerated SMILES a further improvement to a correlation coefficient of 0.68 and a RMS of 0.52 was found.
References in corpus (1)
Cited by in corpus (32)
- Deep learning for molecular design - a review of the state of the art
- Learning Deep Generative Models of Graphs
- Syntax-Directed Variational Autoencoder for Structured Data
- Improving Chemical Autoencoder Latent Space and Molecular De novo Generation Diversity with Heteroencoders
- Exploring Chemical Space using Natural Language Processing Methodologies for Drug Discovery
- SMILES Transformer: Pre-trained Molecular Fingerprint for Low Data Drug Discovery
- Towards Explainable Anticancer Compound Sensitivity Prediction via Multimodal Attention-based Convolutional Encoders
- Regression Transformer: Concurrent sequence regression and generation for molecular language modeling
- Generative Models for Automatic Chemical Design
- SMILES2Vec: An Interpretable General-Purpose Deep Neural Network for Predicting Chemical Properties
- Molecular Generation with Recurrent Neural Networks (RNNs)
- Toward automatic generation of control structures for process flow diagrams with large language models
- In silico generation of novel, drug-like chemical matter using the LSTM neural network
- PaccMann: Prediction of anticancer compound sensitivity with multi-modal attention-based neural networks
- Improving Deep Learning Models via Constraint-Based Domain Knowledge: a Brief Survey
- Faster and more diverse de novo molecular optimization with double-loop reinforcement learning using augmented SMILES
- Multimodal Deep Neural Networks using Both Engineered and Learned Representations for Biodegradability Prediction
- A Hitchhiker's Guide to Deep Chemical Language Processing for Bioactivity Prediction
- Exploring Bias in GAN-based Data Augmentation for Small Samples
- Permutation invariant graph-to-sequence model for template-free retrosynthesis and reaction prediction
- Data augmentation for machine learning of chemical process flowsheets
- "Found in Translation": Predicting Outcomes of Complex Organic Chemistry Reactions using Neural Sequence-to-Sequence Models
- Artificial Intelligence based Autonomous Molecular Design for Medical Therapeutic: A Perspective
- High throughput screening with machine learning
- Mixup-breakdown: a consistency training method for improving generalization of speech separation models
- Fréchet ChemNet Distance: A metric for generative models for molecules in drug discovery
- Benchmarking Deep Graph Generative Models for Optimizing New Drug Molecules for COVID-19
- Retrosynthetic reaction prediction using neural sequence-to-sequence models
- PaccMann on SARS-CoV-2: Designing antiviral candidates with conditional generative models
- Evening the Score: Targeting SARS-CoV-2 Protease Inhibition in Graph Generative Models for Therapeutic Candidates
- Phenotypic Profiling of High Throughput Imaging Screens with Generic Deep Convolutional Features
- Analysis of training and seed bias in small molecules generated with a conditional graph-based variational autoencoder -- Insights for practical AI-driven molecule generation