Scalable Training of Trustworthy and Energy-Efficient Predictive Graph Foundation Models for Atomistic Materials Modeling: A Case Study with HydraGNN
arXiv:2406.12909 · doi:10.1007/s11227-025-07029-9
Abstract
We present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multi-task learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art United States Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2,000 GPUs on Perlmutter and 16,000 GPUs on Frontier.
51 pages, 32 figures
References in corpus (21)
- Crystal Graph Convolutional Neural Networks for an Accurate and Interpretable Prediction of Material Properties
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- The Open Catalyst 2020 (OC20) Dataset and Community Challenges
- By-passing the Kohn-Sham equations with machine learning
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property Prediction
- The Open Catalyst 2022 (OC22) Dataset and Challenges for Oxide Electrocatalysts
- polyBERT: A chemical language model to enable fully machine-driven ultrafast polymer informatics
- QM7-X: A comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules
- Deep Neural Network Computes Electron Densities and Energies of a Large Set of Organic Molecules Faster than Density Functional Theory (DFT)
- Scientific Large Language Models: A Survey on Biological & Chemical Domains
- Multi-task graph neural networks for simultaneous prediction of global and atomic properties in ferromagnetic systems
- Fast and stable deep-learning predictions of material properties for solid solution alloys
- ClimateBert: A Pretrained Language Model for Climate-Related Text
- Towards Foundation Models for Materials Science: The Open MatSci ML Toolkit
- OpenGlue: Open Source Graph Neural Net Based Pipeline for Image Matching
- On the Scalability of GNNs for Molecular Graphs
- A scalable constructive algorithm for the optimization of neural network architectures
- MuyGPs: Scalable Gaussian Process Hyperparameter Estimation Using Local Cross-Validation
- MatSciML: A Broad, Multi-Task Benchmark for Solid-State Materials Modeling
- Tune As You Scale: Hyperparameter Optimization For Compute Efficient Training