papers

Publications (12)

cs.DC2025

MOFA: Discovering Materials for Carbon Capture with a GenAI- and Simulation-Based Workflow

Xiaoli Yan, Nathaniel Hudson, Hyun Park +15

We present MOFA, an open-source generative AI (GenAI) plus simulation workflow for high-throughput generation of metal-organic frameworks (MOFs) on large-scale high-performance com…

cs.DC2024

TaPS: A Performance Evaluation Suite for Task-based Execution Frameworks

J. Gregory Pauloski, Valerie Hayot-Sasson, Maxime Gonthier +5

Task-based execution frameworks, such as parallel programming libraries, computational workflow systems, and function-as-a-service platforms, enable the composition of distinct tas…

cs.DC2020

Reliable Broadcast in Practical Networks: Algorithm and Evaluation

Yingjian Wu, Haochen Pan, Saptaparni Kumar +1

Reliable broadcast is an important primitive to ensure that a source node can reliably disseminate a message to all the non-faulty nodes in an asynchronous and failure-prone networ…

cs.DC2025

Optimizing Fine-Grained Parallelism Through Dynamic Load Balancing on Multi-Socket Many-Core Systems

Wenyi Wang, Maxime Gonthier, Poornima Nookala +4

Achieving efficient task parallelism on many-core architectures is an important challenge. The widely used GNU OpenMP implementation of the popular OpenMP parallel programming mode…

cs.DC2021

Rabia: Simplifying State-Machine Replication Through Randomization

Haochen Pan, Jesse Tuglu, Neo Zhou +6

We introduce Rabia, a simple and high performance framework for implementing state-machine replication (SMR) within a datacenter. The main innovation of Rabia is in using randomiza…

astro-ph.HE2025

RADAR-Radio Afterglow Detection and AI-driven Response: A Federated Framework for Gravitational Wave Event Follow-Up

Parth Patel, Alessandra Corsi, E. A. Huerta +11

The landmark detection of both gravitational waves (GWs) and electromagnetic (EM) radiation from the binary neutron star merger GW170817 has spurred efforts to streamline the follo…

cs.DC2025

Experiences with Model Context Protocol Servers for Science and High Performance Computing

Haochen Pan, Ryan Chard, Reid Mello +12

Large language model (LLM)-powered agents are increasingly used to plan and execute scientific workflows, yet most research cyberinfrastructure (CI) exposes heterogeneous APIs and…

cs.DC2025

WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks

Sicheng Zhou, Zhuozhao Li, Valérie Hayot-Sasson +6

Failures in Task-based Parallel Programming (TBPP) can severely degrade performance and result in incomplete or incorrect outcomes. Existing failure-handling approaches, including…

cs.DC2025

D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage

Maxime Gonthier, Dante D. Sanchez-Gallegos, Haochen Pan +8

The exponential growth of data necessitates distributed storage models, such as peer-to-peer systems and data federations. While distributed storage can reduce costs and increase r…

cs.DC2026

Icicle: Scalable Metadata Indexing and Real-Time Monitoring for HPC File Systems

Haochen Pan, Ryan Chard, Song Young Oh +7

Modern HPC file systems can contain billions of files and hundreds of petabytes of data, making even simple questions increasingly intractable to answer. Traditional file system ut…

cs.DC2024

Octopus: Experiences with a Hybrid Event-Driven Architecture for Distributed Scientific Computing

Haochen Pan, Ryan Chard, Sicheng Zhou +7

Scientific research increasingly relies on distributed computational resources, storage systems, networks, and instruments, ranging from HPC and cloud systems to edge devices. Even…

cs.DC2025

DynoStore: A wide-area distribution system for the management of data over heterogeneous storage

Dante D. Sanchez-Gallegos, J. L. Gonzalez-Compean, Maxime Gonthier +6

Data distribution across different facilities offers benefits such as enhanced resource utilization, increased resilience through replication, and improved performance by processin…