activity
20242026
collaborators

7 papers

cs.DC2026

Icicle: Scalable Metadata Indexing and Real-Time Monitoring for HPC File Systems

Haochen Pan, Ryan Chard, Song Young Oh +7

Modern HPC file systems can contain billions of files and hundreds of petabytes of data, making even simple questions increasingly intractable to answer. Traditional file system ut…

cs.MA2026

Empowering Scientific Workflows with Federated Agents

Alok Kamatar, J. Gregory Pauloski, Yadu Babuji +5

Agentic systems, in which diverse agents cooperate to tackle challenging problems, are exploding in popularity in the AI community. However, existing agentic frameworks take a rela…

cs.DC2025

FIRST: Federated Inference Resource Scheduling Toolkit for Scientific AI Model Access

Aditya Tanikanti, Benoit Côté, Yanfei Guo +9

We present the Federated Inference Resource Scheduling Toolkit (FIRST), a framework enabling Inference-as-a-Service across distributed High-Performance Computing (HPC) clusters. FI…

cs.DC2025

Experiences with Model Context Protocol Servers for Science and High Performance Computing

Haochen Pan, Ryan Chard, Reid Mello +12

Large language model (LLM)-powered agents are increasingly used to plan and execute scientific workflows, yet most research cyberinfrastructure (CI) exposes heterogeneous APIs and…

astro-ph.HE2025

RADAR-Radio Afterglow Detection and AI-driven Response: A Federated Framework for Gravitational Wave Event Follow-Up

Parth Patel, Alessandra Corsi, E. A. Huerta +11

The landmark detection of both gravitational waves (GWs) and electromagnetic (EM) radiation from the binary neutron star merger GW170817 has spurred efforts to streamline the follo…

cs.DC2025

WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks

Sicheng Zhou, Zhuozhao Li, Valérie Hayot-Sasson +6

Failures in Task-based Parallel Programming (TBPP) can severely degrade performance and result in incomplete or incorrect outcomes. Existing failure-handling approaches, including…