7 papers
Machine Learning-Driven Predictive Resource Management in Complex Science Workflows
Tasnuva Chowdhury, Tadashi Maeno, Fatih Furkan Akman +23
The collaborative efforts of large communities in science experiments, often comprising thousands of global members, reflect a monumental commitment to exploration and discovery. R…
Data Management System Analysis for Distributed Computing Workloads
Kuan-Chieh Hsu, Sairam Sri Vatsavai, Ozgur O. Kilic +20
Large-scale international collaborations such as ATLAS rely on globally distributed workflows and data management to process, move, and store vast volumes of data. ATLAS's Producti…
Towards an Introspective Dynamic Model of Globally Distributed Computing Infrastructures
Ozgur O. Kilic, David K. Park, Yihui Ren +18
Large-scale scientific collaborations like ATLAS, Belle II, CMS, DUNE, and others involve hundreds of research institutes and thousands of researchers spread across the globe. Thes…
Few-shot Personalization of LLMs with Mis-aligned Responses
Jaehyung Kim, Yiming Yang
As the diversity of users increases, the capability of providing personalized responses by large language models (LLMs) has become increasingly important. Existing approaches have…
Alternative Mixed Integer Linear Programming Optimization for Joint Job Scheduling and Data Allocation in Grid Computing
Shengyu Feng, Jaehyung Kim, Yiming Yang +18
This paper presents a novel approach to the joint optimization of job scheduling and data allocation in grid computing environments. We formulate this joint optimization problem as…
AI Surrogate Model for Distributed Computing Workloads
David K. Park, Yihui Ren, Ozgur O. Kilic +18
Large-scale international scientific collaborations, such as ATLAS, Belle II, CMS, and DUNE, generate vast volumes of data. These experiments necessitate substantial computational…