Linking the Neural Machine Translation and the Prediction of Organic Chemistry Reactions
arXiv:1612.09529
Abstract
Finding the main product of a chemical reaction is one of the important problems of organic chemistry. This paper describes a method of applying a neural machine translation model to the prediction of organic chemical reactions. In order to translate 'reactants and reagents' to 'products', a gated recurrent unit based sequence-to-sequence model and a parser to generate input tokens for model from reaction SMILES strings were built. Training sets are composed of reactions from the patent databases, and reactions manually generated applying the elementary reactions in an organic chemistry textbook of Wade. The trained models were tested by examples and problems in the textbook. The prediction process does not need manual encoding of rules (e.g., SMARTS transformations) to predict products, hence it only needs sufficient training reaction sets to learn new types of reactions.
19 pages, 5 figures
References in corpus (2)
Cited by in corpus (8)
- Combining Machine Learning and Computational Chemistry for Predictive Insights Into Chemical Systems
- Exploring Chemical Space using Natural Language Processing Methodologies for Drug Discovery
- Root-aligned SMILES: A Tight Representation for Chemical Reaction Prediction
- Graph Transformation Policy Network for Chemical Reaction Prediction
- Permutation invariant graph-to-sequence model for template-free retrosynthesis and reaction prediction
- "Found in Translation": Predicting Outcomes of Complex Organic Chemistry Reactions using Neural Sequence-to-Sequence Models
- Modern Hopfield Networks for Few- and Zero-Shot Reaction Template Prediction
- Retrosynthetic reaction prediction using neural sequence-to-sequence models