11 papers
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
Tasnuva Chowdhury, Tadashi Maeno, Fatih Furkan Akman +23
The collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. R…
JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows
Vladislav Esaulov, Jieyang Chen, Norbert Podhorszki +4
In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise…
Data Management System Analysis for Distributed Computing Workloads
Kuan-Chieh Hsu, Sairam Sri Vatsavai, Ozgur O. Kilic +20
Large-scale international collaborations such as ATLAS rely on globally distributed workflows and data management to process, move, and store vast volumes of data. ATLAS's Producti…
CGSim: A Simulation Framework for Large Scale Distributed Computing Environment
Sairam Sri Vatsavai, Raees Khan, Kuan-Chieh Hsu +20
Large-scale distributed computing infrastructures such as the Worldwide LHC Computing Grid (WLCG) require comprehensive simulation tools for evaluating performance, testing new alg…
The Artificial Scientist -- in-transit Machine Learning of Plasma Simulations
Jeffrey Kelling, Vicente Bolea, Michael Bussmann +19
Increasing HPC cluster sizes and large-scale simulations that produce petabytes of data per run, create massive IO and storage challenges for analysis. Deep learning-based techniqu…
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
Ozgur O. Kilic, David K. Park, Yihui Ren +18
Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. Thes…