collaborators

10 papers

cs.AI2026

Automated Data Readiness for Scientific AI

Sean R. Wilkinson, Valentine G. Anantharaj, Jong Youl Choi +8

Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as AI training data. However, no existing f…

cs.DC2025

Designing FAIR Workflows at OLCF: Building Scalable and Reusable Ecosystems for HPC Science

Sean R. Wilkinson, Patrick Widener, Sarp Oral +1

High Performance Computing (HPC) centers provide advanced infrastructure that enables scientific research at extreme scale. These centers operate with hardware configurations, soft…

cs.LG2025

Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training

Wesley Brewer, Murali Meena Gopalakrishnan, Matthias Maiterth +12

With the end of Moore's law and Dennard scaling, efficient training increasingly requires rethinking data volume. Can we train better models with significantly less data via intell…

cs.DC2025

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

Anjus George, Michael Brim, Christopher Zimmer +3

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine th…

cs.DC2025

A Study on Messaging Trade-offs in Data Streaming for Scientific Workflows

Anjus George, Michael J. Brim, Christopher Zimmer +4

Memory-to-memory data streaming is essential for modern scientific workflows that require near real-time data analysis, experimental steering, and informed decision-making during e…

cs.LG2025

Scaling Up Data Parallelism in Decentralized Deep Learning

Bing Xie, Junqi Yin, Zhenyu Zhou +2

Although it has been extensively explored in theory, decentralized learning is not yet green-lighted for production use, largely due to a lack of stability, scalability, and genera…