Modeling Industrial ADMET Data with Multitask Networks
arXiv:1606.08793
Abstract
Deep learning methods such as multitask neural networks have recently been applied to ligand-based virtual screening and other drug discovery applications. Using a set of industrial ADMET datasets, we compare neural networks to standard baseline models and analyze multitask learning effects with both random cross-validation and a more relevant temporal validation scheme. We confirm that multitask learning can provide modest benefits over single-task models and show that smaller datasets tend to benefit more than larger datasets from multitask learning. Additionally, we find that adding massive amounts of side information is not guaranteed to improve performance relative to simpler multitask learning. Our results emphasize that multitask effects are highly dataset-dependent, suggesting the use of dataset-specific models to maximize overall performance.
See "Version information" section
Cited by in corpus (9)
- Deep Learning in Pharmacogenomics: From Gene Regulation to Patient Stratification
- Chemception: A Deep Neural Network with Minimal Chemistry Knowledge Matches the Performance of Expert-developed QSAR/QSPR Models
- Generating Focussed Molecule Libraries for Drug Discovery with Recurrent Neural Networks
- MoleculeNet: A Benchmark for Molecular Machine Learning
- Deep Learning for Computational Chemistry
- Chemi-net: a graph convolutional network for accurate drug property prediction
- Benchmarking Accuracy and Generalizability of Four Graph Neural Networks Using Large In Vitro ADME Datasets from Different Chemical Spaces
- Meta-Learning GNN Initializations for Low-Resource Molecular Property Prediction
- MolDesigner: Interactive Design of Efficacious Drugs with Deep Learning