most citedAutomated Data Readiness for Scientific AI

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.DC2026

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji +5

Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations,…

cs.DL2026

SetGo: Metadata Readiness for Scientific AI Datasets

Sean R. Wilkinson, Polina Shpilker, Wesley Brewer

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for D…

cs.AI20261 cited

Automated Data Readiness for Scientific AI

Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi +8

Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing f…

cs.LG2025

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

Wesley Brewer, Murali Meena Gopalakrishnan, Matthias Maiterth +12

With the end of Moore's law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intell…

cs.DC2025

Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies

Wesley Brewer, Matthias Maiterth, Damien Fay

The rapid growth of AI supercomputing is creating unprecedented power demands, with next-generation GPU datacenters requiring hundreds of megawatts and producing fast, large swings…

cs.DC2025

HPC Digital Twins for Evaluating Scheduling Policies, Incentive Structures and their Impact on Power and Cooling

Matthias Maiterth, Wesley H. Brewer, Jaya S. Kuruvella +8

Schedulers are critical for optimal resource utilization in high-performance computing. Traditional methods to evaluate schedulers are limited to post-deployment analysis, or simul…