Exploring Chemical Space using Natural Language Processing Methodologies for Drug Discovery
arXiv:2002.06053 · doi:10.1016/j.drudis.2020.01.020
Abstract
Text-based representations of chemicals and proteins can be thought of as unstructured languages codified by humans to describe domain-specific knowledge. Advances in natural language processing (NLP) methodologies in the processing of spoken languages accelerated the application of NLP to elucidate hidden knowledge in textual representations of these biochemical entities and then use it to construct models to predict molecular properties or to design novel molecules. This review outlines the impact made by these advances on drug discovery and aims to further the dialogue between medicinal chemists and computer scientists.
References in corpus (11)
- Sequence to Sequence Learning with Neural Networks
- From Frequency to Meaning: Vector Space Models of Semantics
- Deep learning for molecular design - a review of the state of the art
- What do we need to build explainable AI systems for the medical domain?
- Predicting Organic Reaction Outcomes with Weisfeiler-Lehman Network
- WideDTA: prediction of drug-target binding affinity
- Linking the Neural Machine Translation and the Prediction of Organic Chemistry Reactions
- Application of generative autoencoder in de novo molecular design
- Topic-Guided Variational Autoencoders for Text Generation
- Synergy Effect between Convolutional Neural Networks and the Multiplicity of SMILES for Improvement of Molecular Prediction
- In silico generation of novel, drug-like chemical matter using the LSTM neural network