Probing out-of-distribution generalization in machine learning for materials
arXiv:2406.06489 · doi:10.1038/s43246-024-00731-w
Abstract
Scientific machine learning (ML) endeavors to develop generalizable models with broad applicability. However, the assessment of generalizability is often based on heuristics. Here, we demonstrate in the materials science setting that heuristics based evaluations lead to substantially biased conclusions of ML generalizability and benefits of neural scaling. We evaluate generalization performance in over 700 out-of-distribution tasks that features new chemistry or structural symmetry not present in the training data. Surprisingly, good performance is found in most tasks and across various ML models including simple boosted trees. Analysis of the materials representation space reveals that most tasks contain test data that lie in regions well covered by training data, while poorly-performing tasks contain mainly test data outside the training domain. For the latter case, increasing training set size or training time has marginal or even adverse effects on the generalization performance, contrary to what the neural scaling paradigm assumes. Our findings show that most heuristically-defined out-of-distribution tests are not genuinely difficult and evaluate only the ability to interpolate. Evaluating on such tasks rather than the truly challenging ones can lead to an overestimation of generalizability and benefits of scaling.
References in corpus (15)
- A General-Purpose Machine Learning Framework for Predicting Properties of Inorganic Materials
- A Universal Graph Deep Learning Interatomic Potential for the Periodic Table
- The Open Catalyst 2020 (OC20) Dataset and Community Challenges
- Atomistic Line Graph Neural Network for Improved Materials Property Predictions
- The Joint Automated Repository for Various Integrated Simulations (JARVIS) for data-driven materials design
- On scientific understanding with artificial intelligence
- The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts
- Towards self-driving laboratories: The central role of density functional theory in the AI age
- A map of single-phase high-entropy alloys
- A critical examination of robustness and generalizability of machine learning prediction of materials properties
- On the redundancy in large material datasets: efficient and robust learning with less data
- Recent progress in the JARVIS infrastructure for next-generation data-driven materials design
- ET-AL: Entropy-Targeted Active Learning for Bias Mitigation in Materials Data
- A Deep-learning Model for Fast Prediction of Vacancy Formation in Diverse Materials
- Efficient first principles based modeling via machine learning: from simple representations to high entropy materials