4 papers
Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training
Wesley Brewer, Murali Meena Gopalakrishnan, Matthias Maiterth +12
With the end of Moore's law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intell…
Scaling Up Data Parallelism in Decentralized Deep Learning
Bing Xie, Junqi Yin, Zhenyu Zhou +2
Although it has been extensively explored in theory, decentralized learning is not yet green-lighted for production use, largely due to a lack of stability, scalability, and genera…
HPC Digital Twins for Evaluating Scheduling Policies, Incentive Structures and their Impact on Power and Cooling
Matthias Maiterth, Wesley H. Brewer, Jaya S. Kuruvella +8
Schedulers are critical for optimal resource utilization in high-performance computing. Traditional methods to evaluate schedulers are limited to post-deployment analysis, or simul…
Data Readiness for Scientific AI at Scale
Wesley Brewer, Patrick Widener, Valentine Anantharaj +4
This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows acros…