papers

Publications (48)

cs.DC2025

RHAPSODY: Execution of Hybrid AI-HPC Workflows at Scale

Aymen Alsaadi, Mason Hooten, Mariya Goliyad +14

Hybrid AI-HPC workflows combine large-scale simulation, training, high-throughput inference, and tightly coupled, agent-driven control within a single execution campaign. These wor…

cs.DC2023

Asynchronous Execution of Heterogeneous Tasks in ML-driven HPC Workflows

Vincent R. Pascuzzi, Ozgur O. Kilic, Matteo Turilli +1

Heterogeneous scientific workflows consist of numerous types of tasks that require executing on heterogeneous resources. Asynchronous execution of those tasks is crucial to improve…

cs.SE2019

Designing Workflow Systems Using Building Blocks

Matteo Turilli, Andre Merzky, Vivek Balasubramanian +2

We suggest there is a need for a fresh perspective on the design and development of workflow systems and argue for a building blocks approach. We outline a description of this appr…

cs.DC2019

Middleware Building Blocks for Workflow Systems

Matteo Turilli, Vivek Balasubramanian, Andre Merzky +2

This paper describes a building blocks approach to the design of scientific workflow systems. We discuss RADICAL-Cybertools as one implementation of the building blocks concept, sh…

cs.DC2022

The Ghost of Performance Reproducibility Past

Srinivasan Ramesh, Mikhail Titov, Matteo Turilli +2

The importance of ensemble computing is well established. However, executing ensembles at scale introduces interesting performance fluctuations that have not been well investigated…

physics.comp-ph2022

Pipeline for Automating Compliance-based Elimination and Extension (PACE2): A Systematic Framework for High-throughput Biomolecular Material Simulation Workflows

Srinivas C. Mushnoori, Ethan Zang, Akash Banerjee +5

The formation of biomolecular materials via dynamical interfacial processes such as self-assembly and fusion, for diverse compositions and external conditions, can be efficiently p…

cs.DC2022

RADICAL-Pilot and Parsl: Executing Heterogeneous Workflows on HPC Platforms

Aymen Alsaadi, Logan Ward, Andre Merzky +4

Workflows applications are becoming increasingly important to support scientific discovery. That is leading to a proliferation of workflow management systems and, thus, to a fragme…

cs.DC2022

RAPTOR: Ravenous Throughput Computing

Andre Merzky, Matteo Turilli, Shantenu Jha

We describe the design, implementation and performance of the RADICAL-Pilot task overlay (RAPTOR). RAPTOR enables the execution of heterogeneous tasks -- i.e., functions and execut…

cs.DC2021

Workflows Community Summit: Bringing the Scientific Workflows Community Together

Rafael Ferreira da Silva, Henri Casanova, Kyle Chard +42

Scientific workflows have been used almost universally across scientific domains, and have underpinned some of the most significant discoveries of the past several decades. Many of…

cs.DC2018

Concurrent and Adaptive Extreme Scale Binding Free Energy Calculations

Jumana Dakka, Kristof Farkas-Pall, Matteo Turilli +3

The efficacy of drug treatments depends on how tightly small molecules bind to their target proteins. The rapid and accurate quantification of the strength of these interactions (a…

cs.DC2020

Comparing Workflow Application Designs for High Resolution Satellite Image Analysis

Aymen Al-Saadi, Ioannis Paraskevakos, Bento Collares Gonçalves +3

Very High Resolution satellite and aerial imagery are used to monitor and conduct large scale surveys of ecological systems. Convolutional Neural Networks have successfully been em…

cs.DC2025

Scalable Runtime Architecture for Data-driven, Hybrid HPC and ML Workflow Applications

Andre Merzky, Mikhail Titov, Matteo Turilli +3

Hybrid workflows combining traditional HPC and novel ML methodologies are transforming scientific computing. This paper presents the architecture and implementation of a scalable r…

cs.DC2018

Synapse: Synthetic Application Profiler and Emulator

Andre Merzky, Ming Tai Ha, Matteo Turilli +1

Motivated by the need to emulate workload execution characteristics on high-performance and distributed heterogeneous resources, we introduce Synapse. Synapse is used as a proxy ap…

cs.DC2025

Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads

Andre Merzky, Mikhail Titov, Matteo Turilli +1

Scientific workflows increasingly involve both HPC and machine-learning tasks, combining MPI-based simulations, training, and inference in a single execution. Launchers such as Slu…

cs.SE2024

Exascale Workflow Applications and Middleware: An ExaWorks Retrospective

Aymen Alsaadi, Mihael Hategan-Marandiuc, Ketan Maheshwari +9

Exascale computers offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific discover…

q-bio.BM2019

Deep Generative Model Driven Protein Folding Simulation

Heng Ma, Debsindhu Bhowmik, Hyungro Lee +4

Significant progress in computer hardware and software have enabled molecular dynamics (MD) simulations to model complex biological phenomena such as protein folding. However, enab…

cs.SE2024

ExaWorks Software Development Kit: A Robust and Scalable Collection of Interoperable Workflow Technologies

Matteo Turilli, Mihael Hategan-Marandiuc, Mikhail Titov +13

Scientific discovery increasingly requires executing heterogeneous scientific workflows on high-performance computing (HPC) platforms. Heterogeneous workflows contain different typ…

cs.DC2021

Pandemic Drugs at Pandemic Speed: Infrastructure for Accelerating COVID-19 Drug Discovery with Hybrid Machine Learning- and Physics-based Simulations on High Performance Computers

Agastya P. Bhati, Shunzhou Wan, Dario Alfè +26

The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottle…

cs.DC2023

Workflows Community Summit 2022: A Roadmap Revolution

Rafael Ferreira da Silva, Rosa M. Badia, Venkat Bala +102

Scientific workflows have become integral tools in broad scientific computing use cases. Science discovery is increasingly dependent on workflows to orchestrate large and complex s…

q-bio.BM2021

Protein-Ligand Docking Surrogate Models: A SARS-CoV-2 Benchmark for Deep Learning Accelerated Virtual Screening

Austin Clyde, Thomas Brettin, Alexander Partin +8

We propose a benchmark to study surrogate model accuracy for protein-ligand docking. We share a dataset consisting of 200 million 3D complex structures and 2D structure scores acro…

cs.DC2024

Workflow Mini-Apps: Portable, Scalable, Tunable & Faithful Representations of Scientific Workflows

Ozgur Ozan Kilic, Tianle Wang, Matteo Turilli +4

Workflows are critical for scientific discovery. However, the sophistication, heterogeneity, and scale of workflows make building, testing, and optimizing them increasingly challen…

cs.DC2020

Scalable HPC and AI Infrastructure for COVID-19 Therapeutics

Hyungro Lee, Andre Merzky, Li Tan +15

COVID-19 has claimed more 1 million lives and resulted in over 40 million infections. There is an urgent need to identify drugs that can inhibit SARS-CoV-2. In response, the DOE re…

cs.DC2020

Workflow Design Analysis for High Resolution Satellite Image Analysis

Ioannis Paraskevakos, Matteo Turilli, Bento Collares Gonçalves +2

Ecological sciences are using imagery from a variety of sources to monitor and survey populations and ecosystems. Very High Resolution (VHR) satellite imagery provide an effective…

cs.DC2025

Adaptive Protein Design Protocols and Middleware

Aymen Alsaadi, Jonathan Ash, Mikhail Titov +4

Computational protein design is experiencing a transformation driven by AI/ML. However, the range of potential protein sequences and structures is astronomically vast, even for mod…

cs.DC2024

Scaling on Frontier: Uncertainty Quantification Workflow Applications using ExaWorks to Enable Full System Utilization

Mikhail Titov, Robert Carson, Matthew Rolchigo +6

When running at scale, modern scientific workflows require middleware to handle allocated resources, distribute computing payloads and guarantee a resilient execution. While indivi…

cs.DC2021

Evaluating Distributed Execution of Workloads

Matteo Turilli, Yadu Nand Babuji, Andre Merzky +4

Resource selection and task placement for distributed execution poses conceptual and implementation difficulties. Although resource selection and task placement are at the core of…

cs.DC2016

Integrating Abstractions to Enhance the Execution of Distributed Applications

Matteo Turilli, Feng Liu, Zhao Zhang +5

One of the factors that limits the scale, performance, and sophistication of distributed applications is the difficulty of concurrently executing them on multiple distributed compu…

cs.DC2024

Design and Implementation of an Analysis Pipeline for Heterogeneous Data

Arup Kumar Sarker, Aymen Alsaadi, Niranda Perera +8

Managing and preparing complex data for deep learning, a prevalent approach in large-scale data science can be challenging. Data transfer for model training also presents difficult…

cs.DC2019

Characterizing the Performance of Executing Many-tasks on Summit

Matteo Turilli, Andre Merzky, Thomas Naughton +2

Many scientific workloads are comprised of many tasks, where each task is an independent simulation or analysis of data. The execution of millions of tasks on heterogeneous HPC pla…

cs.CE2022

A Scalable Solution for Running Ensemble Simulations for Photovoltaic Energy

Weiming Hu, Guido Cervone, Matteo Turilli +2

This chapter proposes and provides an in-depth discussion of a scalable solution for running ensemble simulation for solar energy production. Generating a forecast ensemble is comp…

cs.DC2021

Workflows Community Summit: Advancing the State-of-the-art of Scientific Workflows Management Systems Research and Development

Rafael Ferreira da Silva, Henri Casanova, Kyle Chard +55

Scientific workflows are a cornerstone of modern scientific computing, and they have underpinned some of the most significant discoveries of the last decade. Many of these workflow…

cs.DC2017

High-Throughput Computing on High-Performance Platforms: A Case Study

Danila Oleynik, Sergey Panitkin, Matteo Turilli +6

The computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size resourc…

cs.DC2021

ExaWorks: Workflows for Exascale

Aymen Al-Saadi, Dong H. Ahn, Yadu Babuji +12

Exascale computers will offer transformative capabilities to combine data-driven and learning-based approaches with traditional simulation applications to accelerate scientific dis…

cs.SE2019

RADICAL-Cybertools: Middleware Building Blocks for Scalable Science

Vivek Balasubramanian, Shantenu Jha, Andre Merzky +1

RADICAL-Cybertools (RCT) are a set of software systems that serve as middleware to develop efficient and effective tools for scientific computing. Specifically, RCT enable executin…

cs.DC2024

Hydra: Brokering Cloud and HPC Resources to Support the Execution of Heterogeneous Workloads at Scale

Aymen Alsaadi, Shantenu Jha, Matteo Turilli

Scientific discovery increasingly depends on middleware that enables the execution of heterogeneous workflows on heterogeneous platforms One of the main challenges is to design sof…

cs.DC2022

AI-coupled HPC Workflows

Shantenu Jha, Vincent R. Pascuzzi, Matteo Turilli

Increasingly, scientific discovery requires sophisticated and scalable workflows. Workflows have become the ``new applications,'' wherein multi-scale computing campaigns comprise m…

cs.DC2020

IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEads

Aymen Al Saadi, Dario Alfe, Yadu Babuji +33

The drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2-3 billion to deliver one new drug. This is both too expensive…

cs.DC2016

A Comprehensive Perspective on Pilot-Job Systems

Matteo Turilli, Mark Santcroos, Shantenu Jha

Pilot-Job systems play an important role in supporting distributed scientific computing. They are used to consume more than 700 million CPU hours a year by the Open Science Grid co…

cs.DC2018

Towards General Distributed Resource Selection

Ming Tai Ha, Matteo Turilli, Andre Merzky +1

The advantages of distributing workloads and utilizing multiple distributed resources are now well established. The type and degree of heterogeneity of distributed resources is inc…

cs.DC2018

Harnessing the Power of Many: Extensible Toolkit for Scalable Ensemble Applications

Vivek Balasubramanian, Matteo Turilli, Weiming Hu +5

Many scientific problems require multiple distinct computational tasks to be executed in order to achieve a desired solution. We introduce the Ensemble Toolkit (EnTK) to address th…

cs.DC2022

Coupling streaming AI and HPC ensembles to achieve 100-1000x faster biomolecular simulations

Alexander Brace, Igor Yakushin, Heng Ma +7

Machine learning (ML)-based steering can improve the performance of ensemble-based simulations by allowing for online selection of more scientifically meaningful computations. We p…

cs.CE2019

Adaptive Ensemble Biomolecular Simulations at Scale

Vivek Balasubramanian, Travis Jensen, Matteo Turilli +3

Recent advances in both theory and methods have created opportunities to simulate biomolecular processes more efficiently using adaptive ensemble simulations. Ensemble-based simula…

cs.DC2023

PSI/J: A Portable Interface for Submitting, Monitoring, and Managing Jobs

Mihael Hategan-Marandiuc, Andre Merzky, Nicholson Collier +10

It is generally desirable for high-performance computing (HPC) applications to be portable between HPC systems, for example to make use of more performant hardware, make effective…

cs.DC2018

Using Pilot Systems to Execute Many Task Workloads on Supercomputers

Andre Merzky, Matteo Turilli, Manuel Maldonado +2

High performance computing systems have historically been designed to support applications comprised of mostly monolithic, single-job workloads. Pilot systems decouple workload spe…

cs.DC2021

Design and Performance Characterization of RADICAL-Pilot on Leadership-class Platforms

Andre Merzky, Matteo Turilli, Mikhail Titov +2

Many extreme scale scientific applications have workloads comprised of a large number of individual high-performance tasks. The Pilot abstraction decouples workload specification,…

cs.DC2019

DeepDriveMD: Deep-Learning Driven Adaptive Molecular Simulations for Protein Folding

Hyungro Lee, Heng Ma, Matteo Turilli +3

Simulations of biological macromolecules play an important role in understanding the physical basis of a number of complex processes such as protein folding. Even with increasing c…

cs.DC2018

Design and Performance Characterization of RADICAL-Pilot on Titan

Andre Merzky, Matteo Turilli, Manuel Maldonado +1

Many extreme scale scientific applications have workloads comprised of a large number of individual high-performance tasks. The Pilot abstraction decouples workload specification,…

cs.DC2018

High-throughput Binding Affinity Calculations at Extreme Scales

Jumana Dakka, Matteo Turilli, David W Wright +5

Resistance to chemotherapy and molecularly targeted therapies is a major factor in limiting the effectiveness of cancer treatment. In many cases, resistance can be linked to geneti…