SMILES2Vec: An Interpretable General-Purpose Deep Neural Network for Predicting Chemical Properties
arXiv:1712.02034
Abstract
Chemical databases store information in text representations, and the SMILES format is a universal standard used in many cheminformatics software. Encoded in each SMILES string is structural information that can be used to predict complex chemical properties. In this work, we develop SMILES2vec, a deep RNN that automatically learns features from SMILES to predict chemical properties, without the need for additional explicit feature engineering. Using Bayesian optimization methods to tune the network architecture, we show that an optimized SMILES2vec model can serve as a general-purpose neural network for predicting distinct chemical properties including toxicity, activity, solubility and solvation energy, while also outperforming contemporary MLP neural networks that uses engineered features. Furthermore, we demonstrate proof-of-concept of interpretability by developing an explanation mask that localizes on the most important characters used in making a prediction. When tested on the solubility dataset, it identified specific parts of a chemical that is consistent with established first-principles knowledge with an accuracy of 88%. Our work demonstrates that neural networks can learn technically accurate chemical concept and provide state-of-the-art accuracy, making interpretable deep neural networks a useful tool of relevance to the chemical industry.
Submitted to SIGKDD 2018
References in corpus (10)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Neural Architecture Search with Reinforcement Learning
- ANI-1: An extensible neural network potential with DFT accuracy at force field computational cost
- cuDNN: Efficient Primitives for Deep Learning
- Self-Normalizing Neural Networks
- Massively Multitask Networks for Drug Discovery
- SMILES Enumeration as Data Augmentation for Neural Network Modeling of Molecules
- Multi-task Neural Networks for QSAR Predictions
- Chemception: A Deep Neural Network with Minimal Chemistry Knowledge Matches the Performance of Expert-developed QSAR/QSPR Models
- MoleculeNet: A Benchmark for Molecular Machine Learning
Cited by in corpus (17)
- Autonomous discovery in the chemical sciences part II: Outlook
- Exploring Chemical Space using Natural Language Processing Methodologies for Drug Discovery
- Towards Explainable Anticancer Compound Sensitivity Prediction via Multimodal Attention-based Convolutional Encoders
- Machine learning and AI-based approaches for bioactive ligand discovery and GPCR-ligand recognition
- A novel methodology on distributed representations of proteins using their interacting ligands
- IRNet: A General Purpose Deep Residual Regression Framework for Materials Discovery
- Synergy Effect between Convolutional Neural Networks and the Multiplicity of SMILES for Improvement of Molecular Prediction
- PaccMann: Prediction of anticancer compound sensitivity with multi-modal attention-based neural networks
- CheMixNet: Mixed DNN Architectures for Predicting Chemical Properties using Multiple Molecular Representations
- Multimodal Deep Neural Networks using Both Engineered and Learned Representations for Biodegradability Prediction
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Artificial Intelligence in Drug Discovery: Applications and Techniques
- End-to-End Attention-based Image Captioning
- Probabilistic Generative Deep Learning for Molecular Design
- Automated and Explainable Ontology Extension Based on Deep Learning: A Case Study in the Chemical Domain
- Explanatory Masks for Neural Network Interpretability
- IL-Net: Using Expert Knowledge to Guide the Design of Furcated Neural Networks