Understanding molecular representations in machine learning: The role of uniqueness and target similarity
arXiv:1608.06194 · doi:10.1063/1.4964627
Abstract
The predictive accuracy of Machine Learning (ML) models of molecular properties depends on the choice of the molecular representation. Based on the postulates of quantum mechanics, we introduce a hierarchy of representations which meet uniqueness and target similarity criteria. To systematically control target similarity, we rely on interatomic many body expansions, as implemented in universal force-fields, including Bonding, Angular, and higher order terms (BA). Addition of higher order contributions systematically increases similarity to the true potential energy and predictive accuracy of the resulting ML models. We report numerical evidence for the performance of BAML models trained on molecular properties pre-calculated at electron-correlated and density functional theory level of theory for thousands of small organic molecules. Properties studied include enthalpies and free energies of atomization, heatcapacity, zero-point vibrational energies, dipole-moment, polarizability, HOMO/LUMO energies and gap, ionization potential, electron affinity, and electronic excitations. After training, BAML predicts energies or electronic properties of out-of-sample molecules with unprecedented accuracy and speed.
References in corpus (1)
Cited by in corpus (33)
- Machine Learning Force Fields
- Molecular Contrastive Learning of Representations via Graph Neural Networks
- Machine Learning Unifies the Modelling of Materials and Molecules
- Physics-inspired structural representations for molecules and materials
- Opportunities and Challenges for Machine Learning in Materials Science
- Machine learning for electronically excited states of molecules
- ANI-1: A data set of 20M off-equilibrium DFT calculations for organic molecules
- Resolving transition metal chemical space: feature selection for machine learning and structure-property relationships
- Towards self-driving laboratories: The central role of density functional theory in the AI age
- Unsupervised machine learning in atomistic simulations, between predictions and understanding
- Neural Network Potentials for Chemistry: Concepts, Applications and Prospects
- Machine learning based energy-free structure predictions of molecules (closed and open-shell), transition states, and solids
- Maximizing information from chemical engineering data sets: Applications to machine learning
- PiNN: A Python Library for Building Atomic Neural Networks of Molecules and Materials
- Recursive evaluation and iterative contraction of -body equivariant features
- Deep Learning for UV Absorption Spectra with SchNarc: First Steps Towards Transferability in Chemical Compound Space
- Using Molecular Embeddings in QSAR Modeling: Does it Make a Difference?
- Global optimization of atomic structures with gradient-enhanced Gaussian process regression
- Towards DMC accuracy across chemical space with scalable -QML
- An orbital-based representation for accurate Quantum Machine Learning
- Optimal radial basis for density-based atomic representations
- Kernel based quantum machine learning at record rate : Many-body distribution functionals as compact representations
- SPAM: the Spectrum of Approximated Hamiltonian Matrices representations
- Quantum chemical roots of machine-learning molecular similarity descriptors
- Ab initio machine learning of phase space averages
- Decomposing Chemical Space: Applications to the Machine Learning of Atomic Energies
- Mean-Field Density Matrix Decompositions
- Supervised learning of few dirty bosons with variable particle number
- Generating candidates in global optimization algorithms using complementary energy landscapes
- A smooth basis for atomistic machine learning
- Comment on "Manifolds of quasi-constant SOAP and ACSF fingerprints and the resulting failure to machine learn four body interactions"
- Electronic Descriptors for Supervised Spectroscopic Predictions
- Molecular Hessian matrices from a machine learning random forest regression algorithm